Your Digital Twin. It Already Exists. You Just Didn't Know

How many versions of you exist right now? You think — one. Yourself. A living person with thoughts, habits, a history. One and unrepeatable. In fact — dozens. Maybe hundreds. In databases all over the world, copies of you live right now. They do not know what you are thinking this morning. But they know something about you that you yourself may not realise. They predict your behaviour. They influence decisions that are made about you. They exist independently of you — and you cannot switch them off. This is called a digital twin. And before you say «this sounds like science fiction» — let me show you exactly how it was created. Step by step. Out of things you did every day without thinking.

How your twin is built. From the very beginning. Imagine that someone invisible walks behind you with a notebook. Every day. From the moment you first went online. The first entry in the notebook appeared long ago. Maybe when you registered your first email. Or when you first bought something online. Or when you set up a social-network account. At that moment you entered a name, a date of birth, an address — and this information became the first brick. After that, bricks were added every day.

Every search query. Google stores search history for years. If you never cleared your history and did not use private mode — somewhere there is a record of every question you asked the search engine over the last ten years. What worried you. What you wanted to buy. What you were thinking about at three in the morning when you could not sleep. What illnesses you read about. What news you watched. What angered you enough to search for information about it. This is not just a list of queries. It is a diary of your inner world — only written not by you, but by an algorithm.

Every purchase. Card transaction history is one of the most informative sources of data about a person. Where you buy food — tells about the district you live in and your income level. What exactly you buy — tells about health, habits, marital status, political views. Yes, political views — studies show that from a supermarket shopping list one can predict with high accuracy who a person votes for. When you buy — tells about your daily routine and work schedule. How impulsively — tells about your psychological type. One supermarket receipt is a paragraph in a biography. Years of receipts are a book.

Every movement in space. If you have a smartphone with geolocation on — and most people have it on constantly — somewhere there is a record of every place you have been. Not just «was in the city centre». But exact coordinates accurate to within a few metres. How often you go to the same place. How long you stay there. At what time. With whom — because if two phones regularly end up in the same place at the same time, the system draws a conclusion about a connection between their owners. Geolocation over a year is a map of your life. Work, home, the doctors you visit, the places where you spend time, the routes you take. All of this says more about you than any questionnaire.

Every interaction on social networks. Not only what you post. How you react to others’ content is data. How long you look at a specific post before scrolling on is data. Whether you stop on photos of people or on texts is data. Which topics make you linger longer is data. The algorithm sees not what you decided to show. It sees what you did not decide to show — involuntary reactions, patterns of attention, unconscious preferences. Facebook internally calls this «passive data» as opposed to the «active» data you consciously enter. Passive data is more accurate. Because you do not control it.

And all of this is only the visible part.

Layers you do not see. There is data you never gave consciously. But which nonetheless exists in your profile. Voice. If you have ever used a voice assistant — Siri, Google Assistant, Alexa — your voice was recorded. Not only the commands. Background conversations that the system «accidentally» captured while waiting for the keyword. Amazon publicly admitted that employees listened to users’ recordings to «improve the quality of recognition». From a voice one can extract more than it seems — emotional state, stress level, signs of certain illnesses, accent and origin. Face. If you have ever uploaded photographs to social networks — and most people have uploaded thousands — the geometry of your face exists in databases. The distance between the eyes, the shape of the cheekbones, the contour of the jaw — this is a biometric identifier that does not change with age and cannot be replaced. Face-recognition systems work in airports, shopping centres, on city streets. In some countries — openly. In others — without a public announcement. Sleep and health patterns. Smart watches, fitness trackers, sleep apps. Pulse, blood pressure, activity level, sleep quality. This data goes to companies’ servers. Insurance companies in a number of countries already offer discounts in exchange for access to fitness-tracker data. A discount today — a health profile forever.

The psychological portrait. This is the least obvious and the most important. From the whole body of data, algorithms extract what you yourself do not know about yourself — or would not want to admit. A 2013 Stanford University study showed that from 68 Facebook likes an algorithm predicts a person’s sexual orientation more accurately than their friends. From 150 likes — more accurately than their parents. From 300 likes — more accurately than their spouse. Think about this for a second. A person who has known you for twenty years and lives with you under one roof — knows you worse than an algorithm that never met you but saw three hundred of your likes. This is not magic. It is statistics. Billions of people left data. The algorithm found patterns. Now it applies them to you. Your twin knows — whether you are prone to depression. How anxious you are. Whether you are impulsive in decisions. How susceptible you are to social pressure. How you react to stress. What motivates you and what paralyses you. You told no one this. You simply lived — liked, searched, bought, walked around with a smartphone in your pocket. That turned out to be enough.

How accurate your twin is. This is where it gets truly interesting. Most people think — all right, the algorithm knows something. But it is approximate. Imprecise. It can be wrong. Yes. It can. But more rarely than you think. And the accuracy grows every year. In 2012 there was a story that became a classic example in conversations about data. The Target retail chain in the US developed an algorithm for predicting pregnancy — on the basis of changes in purchasing behaviour. Women in early pregnancy begin to buy certain things — unscented soap, certain vitamins, large bags. The algorithm learned to recognise this before the woman herself told anyone about the pregnancy. One day an angry man came to the store. His daughter — a high-schooler — had been sent coupons for baby products and maternity clothes by the store. He demanded an explanation — deciding that the store was encouraging his daughter to become pregnant. A few weeks later he called back. He apologised. It turned out the daughter really was pregnant. He found out from the store before he found out from her. The algorithm knew before the father. This is 2012. More than ten years have passed since. There is thousands of times more data. The algorithms have become thousands of times smarter. Your twin today is not an approximate sketch. It is a detailed model. It predicts not only what you will buy next month. It predicts how you will vote. How you will react to stress. Whether you are prone to addictions. When you are most vulnerable to certain messages. And someone uses this. Already now.

The moment that changes one’s attitude to this. There is one thing about the digital twin that I want you to truly understand. Not as a fact — but as a feeling. Your twin does not die with you. The data you left — it does not go anywhere after your death. It continues to exist in databases. Continues to be processed. Continues to be sold. Moreover — there already exist companies that offer to create a «digital continuation» of a person after their death. On the basis of everything they left on the internet — correspondence, voice messages, videos, posts — the algorithm creates a model that can respond to messages, imitate the manner of speech, reproduce the manner of communication. This is no longer fiction. The company HereAfter AI, StoryFile, many others — they work. People pay for a subscription to «talk» with deceased relatives. Who controls this copy? Who does it belong to? What can be done with it? Can it give «consent» on a person’s behalf? Can it be used in advertising? In politics? Legislation cannot keep up. GDPR contains some provisions about the data of the deceased — but they are minimal and interpreted differently in different countries. Your twin will outlive you. And what happens to it next — no one yet knows for sure.

What this means for you today. I am not telling you this to frighten you. I am telling you this because understanding changes behaviour. When you know that every search query is a brick in the building of your twin, you begin to treat the search engine differently. When you understand that geolocation builds a map of your life — you begin to treat app permissions differently. When you realise that likes tell more about you than you think — you begin to treat what you do on social networks differently. This does not mean becoming paranoid. It means becoming aware.

Three things you can do right now. First. Go to myaccount.google.com and open the «Data & privacy» section. Look at what is stored there. Search history, geolocation history, YouTube history. You can delete it — and set up automatic deletion after three months. The twin will not disappear completely — but it will become smaller. Second. Go into the smartphone settings and check which apps have constant access to geolocation. Leave only those that genuinely need it — navigation, maps. To all the rest — access only during use or none. Third. Search for yourself. Type your name into Google. Then into Yandex. Then into specialised people-search services — Spokeo, Pipl, Been Verified. Look at what is publicly available. What you see is the visible part of your twin. The part that anyone who wants to can see without any technical effort. This may be unexpected. Your twin already exists. It is already working. It is already influencing decisions that are made about you. The only difference is — whether you know about it. Now you know.

Who owns your digital twin. And what is done with it while you sleep

Last time we examined how your digital twin is built. What it consists of. How accurate it is. What it knows about you — sometimes more than you yourself. The next question. Who does it belong to? The short answer — not to you. The long answer — let us honestly examine exactly who holds your digital copy in their hands. What they do with it. And what decisions about the living you they make on the basis of the digital you — without your participation, without your knowledge, sometimes against your interests. Because your twin has several owners. And each has its own goal.

The first owner. Corporations. Let us start with the obvious — but not the superficial. Google, Meta, Amazon, Apple, Microsoft. These five companies know more about you than any other organisation on the planet. Including the state you live in. Including your loved ones. Including you yourself. But it is important to understand — they do not merely store data. They build models. There is a difference between «we have data about you» and «we built a model that simulates you». Google long ago moved from the first to the second. When you open search — the algorithm does not merely search by your query. It predicts what you mean based on who you are. Two people who enter the same query get different results. Because the algorithm knows — this person usually searches for this in a professional context, and that one in a personal one. This one lives in Tallinn, that one in Berlin. This one’s search history indicates conservative views, that one’s — liberal. Search is personalised. Reality is personalised. Two people live in different versions of the internet — built by their twins. This is called a filter bubble. You see the world not as it is. You see the world as the algorithm thinks you want to see it — based on who it thinks you are.

Now about Amazon. Most people think that Amazon knows what they buy. This is too narrow a view. Amazon knows what you buy — and when. What you look at but do not buy — and why, presumably. What you put in the cart and delete. How long you read a product description. Which reviews you study. What you search for by voice through Alexa. Which films you watch on Prime Video — and where you stop, rewind, abandon them. What music you listen to at what time of day. If you have a Ring — a smart doorbell — Amazon knows your daily routine accurate to the minute. When you leave. When you return. Who comes to your home. All this data does not lie in separate folders. It is combined into a single model. Your twin in the Amazon ecosystem is a very detailed portrait of a person in their everyday life. And Amazon sells advertising. Advertisers pay for access to people with certain profiles. Your twin is inventory. Goods on a shelf. Only the shelf is in the cloud, and you never see it.

Now about Meta — Facebook and Instagram. Here there is something most people do not know. Meta builds profiles not only on its users. It builds profiles on people who never registered with any of its services. This is called a shadow profile. How it works. You do not use Facebook. You never registered. But your friends use it. And when they sync contacts — your phone number ends up in the system. When someone mentions you in a post — your name ends up in the system. When you visit a site that has the Facebook pixel — an invisible tracker — your browser fingerprint ends up in the system. Meta knows about you. Even if you never gave it permission for this. Even if you on principle do not use its products. The shadow profile exists. You cannot request it through the standard form — because you have no account. You cannot delete it — because officially it does not exist. It simply is. The European Court found this practice to be a GDPR violation. Meta paid a fine. The practice continues in an altered form.

The second owner. The state. Here begins a conversation that many prefer to avoid. But avoiding it means not understanding the full picture. States want data about citizens. This is not news — it has always been so. Population censuses, tax registers, accounting systems — all these are forms of state data collection. The difference is that before, the state collected what it could. Now — what corporations want, and what the state can, together give a picture that never existed in history. In democratic countries this works through lawful mechanisms. Requests to companies on the basis of court decisions. Data-exchange programmes between government bodies. Disclosure obligations for telecommunications operators. But there is also something else. After the terrorist attacks of 11 September 2001 in the US, a series of laws was passed expanding the powers of the intelligence services. In 2013 Edward Snowden showed the world the scale of what was actually happening. The NSA — the National Security Agency — had direct access to the servers of Google, Facebook, Microsoft, Apple, Yahoo. The programme was called PRISM. The data of billions of people around the world was collected systematically — not by court requests, but automatically. This is not a conspiracy theory. It is a documented fact confirmed by official documents. Snowden received political asylum in Russia. The surveillance programmes were partially reformed. But the idea that citizens’ digital data is inaccessible to the state — was destroyed forever.

Now about another level. China. China is building what is customarily called a social credit system. This is not a single system with one switch — it is a collection of dozens of programmes at the federal and regional level. The essence is the same — citizens’ behaviour in the real world and online affects their digital rating. The rating affects access to services. A low rating — you cannot buy a plane or high-speed train ticket. You cannot get a loan. Children cannot enrol in certain schools. Your name is published in open lists of «untrustworthy citizens». A high rating — priority in hiring for government positions, lower interest rates, simplified administrative procedures. The digital twin here is not just a profile for advertising. It is a tool for managing the behaviour of living people in the real world. In real time. China is the most open example. But the principles applied there are discussed and partially implemented in other contexts around the world. Insurance scorings, credit ratings, hiring algorithms — these are softer versions of the same idea. The digital twin as a management tool is not the future. It is the present, to varying degrees of implementation in different countries.

The third owner. Insurance companies. Let us talk about money. The concrete money you pay — or do not pay — depending on what the insurance company knows about you. The traditional insurance model worked simply. You fill in a questionnaire. You state your age, sex, medical history, place of residence. The company calculates the risk on the basis of statistics by group — people of your age, sex, profession on average fall ill like this, have accidents like this, die like this. On this basis the tariff is calculated. This is an averaged approach. You pay not for yourself — you pay for the average person of your group. The new model works differently. The insurer wants to know not the average person of your group. It wants to know you. Specifically. Individually. And data makes this possible. Medical insurance. The company Vitality — operating in the United Kingdom, Germany, other countries — offers a discount on insurance if you connect a fitness tracker and meet activity targets. The logic is beautiful — you lead a healthy lifestyle, fall ill less, pay less. But the flip side — the company receives a continuous stream of data about your health. Pulse, activity, sleep, geolocation during workouts. This builds a medical profile more accurate than any questionnaire. What does the company do with this profile in five years? In ten? If the data shows alarming patterns — will the tariff rise? Will they refuse to renew the insurance? Will they sell the data to other insurers? There is no answer. Because no one explains this at the moment when you agree to a discount for a pedometer. Car insurance. Telematics programmes — black boxes installed in the car or apps on the phone — record exactly how you drive. Speed, sharpness of braking, time of day when you drive, routes. A careful driver gets a discount. It sounds fair. But data about your routes is data about your life. Where you go. How often. At what time. The insurance company knows that you drive to a certain doctor every two weeks. Knows that you regularly visit a district statistically associated with certain risks. Knows that you sometimes drive at night — and draws conclusions from this. Life insurance. Here things are more serious. A number of American insurance companies already use data from social networks when calculating life-insurance tariffs. Publicly they speak of this cautiously. But the algorithms look — what lifestyle a person demonstrates on Instagram. Whether they do extreme sports. Whether they drink alcohol, judging by the photos. What their social circle is. You thought you were posting a party photo for your friends. The insurance algorithm thought differently.

The fourth owner. Employers. This direction is growing fastest — and is least regulated. Let us start with hiring. Large companies have long used algorithms for the initial screening of candidates. A CV is analysed automatically — keywords, structure, even formatting. But this is only the beginning. Companies buy data from brokers — or use specialised HR platforms that aggregate public data about candidates. LinkedIn, GitHub, social-network posts, participation in professional forums, a history of conference talks. All of this is gathered into a profile even before you have sent your CV. Amazon for several years used a hiring algorithm that analysed CVs and gave candidates a score. In 2018 it emerged that the algorithm systematically lowered women’s ratings. Not because someone programmed this deliberately. Because the algorithm was trained on historical hiring data — and historically technical positions hired predominantly men. The algorithm reproduced the bias. Amazon closed the programme. But the principle remained — in other companies, under other names. Now about the monitoring of already-hired employees. The Covid pandemic accelerated the shift to remote work — and with it the explosive growth of the market for employee-surveillance software. They are named beautifully — «productivity platforms», «team-management tools». The essence is different. Programmes like Hubstaff, Time Doctor, Teramind take screenshots of the screen every few minutes. Record keystrokes and mouse movements. Analyse which apps are open and for how long. Some turn on the camera to check that the employee is really at the computer. Analyse correspondence in corporate messengers for «tone» and «loyalty». According to research — after the start of the pandemic the use of such programmes grew by 50% in the first year alone. Now this is the norm in many companies. An employee works from home. Feels free. Does not know that every five minutes their screen is photographed. That an algorithm analyses their correspondence. That their «productivity index» is calculated automatically and affects the decision on renewing the contract. An employee’s digital twin is a constantly updated report on every hour of their working day.

The fifth owner. The one they do not talk about. There is a category of owners of your digital twin that is rarely included in this conversation. Not corporations and not states. People with bad intentions. Data leaks. Constantly. On a massive scale. Inevitably. By statistics — every adult in a developed country has already been in at least one major data leak. Most likely — in several. They just do not know about it. The site haveibeenpwned.com — a free service — allows you to check whether your email was in known leaks. Most people who check for the first time — discover that their data leaked two, three, five times. What happens to leaked data? It ends up on darknet forums. There it is bought. Aggregated with other leaks — because from five different leaks one can assemble a single very complete profile. Sold on. The buyers are various. Spammers — the most harmless. Phishers — more dangerous, they use personal data to make letters look convincing. Fraudsters engaged in identity theft — the most dangerous. Identity theft in the digital age is not the theft of a wallet. It is the use of your digital twin instead of you. Bank accounts are opened with your data. Loans are taken. Companies are registered. Crimes are committed. A person finds out about this when a court summons comes to them. Or when a bank refuses a loan because of debts they did not take. Or when an employer finds their name in a register of debtors. Their twin lived its own life. Without them. At their expense.

The moment when all this comes together. Let me paint the whole picture. One day in the life of your twin — while you simply live. Morning. You wake up. The smart watch recorded your sleep quality — the data went to the manufacturer and to the insurance company they have a partnership with. You check the phone — geolocation confirmed you are home, that is a brick in the profile for advertisers. Day. You drive to work. The phone records the route — the data goes to a mapping service and advertising platforms. You search for something in Google — the query is added to the profile, the advertising rebuilds. You open a news site — dozens of trackers record exactly what you read and for how long. You write to a colleague in the corporate messenger — the algorithm analyses the tone for engagement and loyalty. Lunch. You pay by card in a café — the transaction goes to the bank, from the bank — to analytics systems that update the purchasing-behaviour profile. You scroll Instagram while waiting for your order — the algorithm records where your gaze stops, what emotions the content provokes, updates the psychographic profile. Evening. You watch a series on streaming — the platform records exactly what, for how long, where you stop, what you rewatch. You make an online purchase — the transaction data updates the profile at a broker. You reply to a voice message — the voice is recorded, processed, added to the biometric profile. Night. You fall asleep. Your twin does not sleep. It is sold at advertising auctions. It is processed by the algorithms of an insurance company that recalculates your tariff. It is compared with the profiles of people you do not know in order to draw conclusions about you. This is not paranoia. This is literally what happens. Every day. With everyone who has a smartphone and the internet.

What to do about this. To disappear completely from the digital world is impossible. And it is not necessary. The goal is not to become invisible — the goal is to understand who owns what and where the border runs. Three directions that really change the picture. First — minimising your footprint. Not a zero footprint — a minimal one. Use a search engine that does not save history — DuckDuckGo, Brave Search. Use a browser with tracker protection — Firefox with privacy settings, Brave. This does not make you anonymous — but it significantly reduces the number of bricks added to the building of your twin every day. Second — separation. Do not use one email for everything. For important things — one address. For registrations on sites — another, temporary one. Services like SimpleLogin or AnonAddy create aliases that forward mail to your real address — the site does not know who you really are. When an alias starts receiving spam — you know who sold your data. Third — checking. Go to haveibeenpwned.com right now. Enter your email. See which leaks it was in. This will take thirty seconds — and will show the real picture of how many times your data has already leaked. Your digital twin exists. It is owned by corporations, states, insurance companies, employers — and sometimes people you definitely would not want to give it to. But now you know who they are. And knowledge is already a different position.

Where all this is heading. Artificial intelligence, digital twins and the question no one asks

We began by talking about how your digital twin is built. Then who it belongs to and what is done with it right now. Now I want to talk about where this is heading. Not in a hundred years. In five. In ten. At a distance that is already visible — if you know where to look. And about the question that in this conversation almost no one asks. Not a technical question. Not a legal one. A human one. Where does data end — and the personality begin? Who decides where this border is? And what happens when it does not exist?

First — where we are now. To understand where the train is heading — you have to understand how fast it is already going. 2012. An algorithm predicts pregnancy from supermarket purchases. This seemed astonishing — we talked about it in the first post. 2016. The Cambridge Analytica algorithm predicts political behaviour and psychological type from social-network likes. Used to influence elections in dozens of countries. 2019. Chinese face-recognition systems identify a person in a crowd in fractions of a second. The social credit system links this to a behavioural profile in real time. 2021. The company Clearview AI assembled a database of 10 billion photographs from open sources — social networks, news sites, public registers. Any person in any photo is identified instantly. The service is sold to police departments and private companies. Most people whose photos are in the database — never gave consent. 2023. New-generation language models — GPT-4, Claude, Gemini — are capable of imitating the speech style of a specific person with frightening accuracy if given enough texts written by that person. Voice models clone a voice from a few seconds of recording. Video models create realistic video with the face of any person. 2024. All of this comes together. Look at the trajectory. Every two-three years — a qualitative leap. Not a quantitative one — a qualitative one. This is not merely «there became more data». It is «data began to do what it could not before». And we are at the beginning of the curve — not in the middle and not at the end.

What an AI twin is. And why it is not the same as a data profile. Until recently a digital twin was passive. It was stored in databases. It was analysed. On its basis decisions were made. But by itself it did nothing. Artificial intelligence changes this fundamentally. An AI twin is not a profile. It is a simulation. A model that does not merely describe you — but reproduces you. Answers like you. Makes decisions like you. Reacts to new situations the way — in the algorithm’s opinion — you would react. This is a qualitatively different thing. Let me explain through a concrete example that is already happening — not in a lab, but in real life. The company Soul Machines from New Zealand creates Digital People for corporate clients. These are AI avatars that look like living people, speak in living voices, conduct a dialogue in real time. They are used by banks for customer service, medical companies for consultations, educational platforms for teaching. For now these are not specific copies of specific people. For now. But the technology that makes this possible — is the same that allows a copy of you to be created. All that is needed is data. Enough data.

How much data is needed? Researchers from MIT in 2023 showed that to create a convincing voice model of a person, three seconds of clean voice recording is enough. Three seconds. A video call. A voice message. A public speech. To create a text model that imitates writing style — a few thousand words is enough. An active social-network user writes that much in a few weeks. To create a behaviour model — a history of actions in apps over a few months. Corporations already have all of this. For years. On billions of people. Technically — nothing prevents your simulation from being created today. The question is not the possibility. The question is whether it is permitted — and who controls what is done with it.

Scenarios that are already happening

Not the future. The present — just not everywhere covered.

Scenario one. Digital immortality. In 2021 the American Joshua Barbeau lost his girlfriend — Jessica — in a car accident. Jessica was 23. He could not come to terms with the loss. He turned to a company that created a chatbot trained on Jessica’s correspondence — thousands of messages she had sent through various messengers. The bot answered in her style. Used her expressions. Recalled events from her life. He talked with the bot every day. The story became public. It provoked a huge discussion. Psychologists were divided — some said it helps to get through grief, others that it pathologically hinders acceptance of the loss. But here is the question that in this discussion almost never sounded. Did Jessica consent to this? Did she ever think that her correspondence would become the basis for a simulation after death? That this simulation would continue to exist — speaking on her behalf, representing her — without her participation, without her control? No. She simply wrote messages to friends. The company that created the bot received money for a subscription. The bereaved received an illusion of continuation. Jessica — received nothing. Because she was already gone. But her data — was there. This is not the only case. It is an industry. HereAfter AI, StoryFile, Eternos, Replika — dozens of companies work in this space. The digital-immortality market is estimated in the billions of dollars. No regulation. No standards. No answer to the question — who owns a person’s digital copy after their death.

Scenario two. Personality deepfake. 2019. The chief executive of a British energy company receives a call from — as he thought — his boss at the German parent company. The voice absolutely convincing. The accent correct. The manner of speech familiar. The boss asks to urgently transfer 220,000 euros to a Hungarian supplier. The director transfers it. The call was generated by AI. The voice copied from public speeches of the real executive. 220,000 euros went to fraudsters. This was one of the first documented cases of a voice deepfake in fraud. Since then there have been thousands of such cases. The losses worldwide — billions of dollars annually. But financial fraud is only one scenario. And not the most terrible. In 2023 in several countries cases were recorded where fraudsters called elderly people — in the voice of their children or grandchildren. «Mum, I’ve been in an accident, I need money urgently, don’t call Dad, he’ll be upset.» The voice — indistinguishable. The emotions — convincing. Elderly people transferred money. For this a few seconds of voice recording are needed. Which exist in voice messages. In videos on social networks. In public speeches. Your voice exists on the internet. That is enough.

Scenario three. Political manipulation of a new level. Cambridge Analytica used data to show people targeted political messages. That was the first version. The second version looks different. In the Slovak elections in 2023, a few days before the vote, an audio recording appeared. On it — the leader of a liberal party discusses with a journalist how to falsify the elections. The voice convincing. The conversation detailed. The recording was fabricated. AI voice generation. But the debunking appeared later — while the recording spread instantly. The party lost votes. Maybe it did not change the election result — no one knows for sure. But a precedent was set. A deepfake of a politician saying what they never said. A deepfake of a leader declaring war. A deepfake of a public figure in a situation that never happened. The technology exists. Is available. Becomes cheaper every month. And here is what is important. To create a deepfake of a specific person, that person’s data is needed — voice, video, text. The more data — the more convincing the result. The more active a person is on the internet — the better they can be copied. Public people — politicians, journalists, activists — are most vulnerable. But even an ordinary person with active social networks — is already a sufficient target for a targeted attack.

Scenario four. Predictive systems. This is the quietest scenario. And perhaps — the most significant in the long term. Predictive systems are algorithms that, on the basis of a digital twin, predict a person’s future behaviour. And make decisions on the basis of this prediction — before the behaviour has happened. It sounds like fiction. It is called — Minority Report, remember that film? They arrest someone for a crime not yet committed. This is already happening. In a softened form — but by the same logic. In the US a number of police departments use algorithmic crime-prediction systems. PredPol, ShotSpotter, Palantir. The systems analyse historical crime data, demographic data, social networks — and predict where and when the next crime is likely. Or — who is likely to commit it. The police increase patrolling in the «predicted» zones. Pay heightened attention to «predicted» people. The problem. If an algorithm is trained on historical data that reflects an existing bias — it reproduces and amplifies this bias. Historically, certain districts were patrolled more — which means more crimes were recorded there. The algorithm sees more crimes in these districts — predicts more crimes in the same places. The police patrol even more. The loop has closed. A person comes under heightened attention not because they did something. But because the algorithm decided that they are likely to do it. In lending the same logic. The algorithm predicts the probability of default — and refuses a loan to a person who has not yet violated anything. On the basis of the data of people similar to them. On the basis of patterns they demonstrate — often not even financial ones. In hiring. In insurance. In access to government services. Decisions about you are made on the basis of a prediction of who you will be — not of who you are.

Now — the question everyone avoids. I promised to pose the question that in this conversation almost never sounds. Here it is. If a digital twin is accurate enough to predict your decisions — accurate enough to imitate your speech and behaviour — convincing enough to deceive people who know you — where does the copy end and the original begin? This is not a philosophical question for a beautiful sound. It is a practical question with practical consequences. If a company created a simulation — can it give consent on your behalf? Sign contracts? Make transactions? If your AI twin said something — are you responsible for it? Even if you did not say it? If your digital copy exists after your death — who does it belong to? Your heirs? The company that created it? The state? If an algorithm decided that you are likely to do something — does this have legal consequences for you? Already now — in lending and insurance — it does. Is that normal? Legislation cannot keep up with these questions. GDPR — a good law, an important law — was written when most of these scenarios did not exist. It regulates data. But the border between data and personality — is blurring. Data is a description of you. A simulation is already something else. Something that acts. That makes decisions. That influences the world on your behalf. For this we so far have neither a word. Nor a law. Nor a consensus.

A story that makes this concrete. There is one case I want to tell. Not because it is the most dramatic. Because it is the most precise. 2023. The actress Scarlett Johansson discovered that a company had created and was actively promoting an AI voice assistant — with a voice that was indistinguishable from hers. Without her consent. Without her knowledge. Without compensation. The company did not hack her phone. Did not steal recordings. It simply trained the model on public materials — films, interviews, public speeches — that were freely available. Johansson demanded that they stop. The company refused — claiming that the voice was «accidentally similar» and that they had violated nothing. This is a public person with resources, lawyers, media influence. She was able to raise a fuss. The story became international news. An ordinary person in such a situation has neither the resources nor the tools. And such cases — with ordinary people, without publicity — happen in their thousands. Your voice is public — in videos you filmed. Your face is public — in photos you posted. Your speech style is public — in posts you wrote. Technically — that is enough. Legally — for now a grey zone. In practice — it is happening already now.

What will happen next. An honest forecast. I will not pretend that I know exactly how this will unfold. No one knows. But the direction — is visible. Regulation will strengthen. The European AI Act — the law on artificial intelligence — took effect in 2024. This is the world’s first comprehensive law regulating AI systems. It prohibits a number of the most dangerous applications — including social-scoring systems in the style of the Chinese model, real-time biometric identification in public places in most cases, manipulative systems influencing behaviour. This is an important step. But the law always catches up with technology — it never gets ahead. Technology will accelerate. The cost of creating a convincing deepfake falls every year. What in 2020 required an expensive studio — is today done on an ordinary computer in a few minutes. What today requires a computer — in a few years will be done on a phone. The border between the real and the simulated will blur. Already now most people cannot reliably distinguish a high-quality deepfake from real video. In a few years — this will become practically impossible without special tools. And here is what we are heading towards. A world in which the question «is it really him?» — applied to voice, video, text — has no obvious answer. A world in which your digital copy can act independently of you. A world in which decisions about you are made on the basis of predictions about you — not on the basis of your real actions. This is not an apocalypse. It is simply a new environment. As the industrial revolution was a new environment for people who lived in the agrarian world. Not all of them perished. But those who understood what was happening — adapted better than those who did not.

What to do with this understanding. Three things. Concrete. First — awareness of your digital footprint. Now that you understand that every piece of data you leave — is potentially building material for a simulation of you — you treat what you publish differently. Not in the sense of «publish nothing». In the sense of — understanding the difference between what you want to exist publicly, and what leaks without your decision. Second — critical perception of digital content. The voice you hear on the phone — is it really that person? The video you see — is it a real recording? Is the text written by a human or a model? These questions become mandatory — not a sign of paranoia but a sign of literacy. A simple rule. If you are asked to do something urgent and unusual — call back on a number you know. Directly. Not on the number they called from. This simple step destroys most voice-fraud schemes. Third — participation in the conversation. The laws that will regulate AI twins, digital immortality, predictive systems — are being written now. Or will be written in the next few years. These laws will be better if there is a conscious conversation in society about these topics. Not a technical one — a human one. About what matters. About where the border is. About what we want to permit — and what not. This conversation begins with people simply knowing what it is about. You now know.

In closing

I began this post with a question. Where does data end and the personality begin? Here is my answer. The personality is not data. The personality is a continuous living process. It is the capacity to change. To surprise oneself. To make decisions that no algorithm predicted. To do what was not expected of you — including by yourself. A digital twin is a cast. A very precise one. Sometimes frighteningly precise. But always — of the past. It knows who you were. It predicts who you will be — on the basis of who you were. But it does not know who you will decide to become. That is the border. Data describes a trajectory. A personality can leave it. And as long as this is so — as long as a person is capable of surprising the algorithm — the conversation about data, rights and digital hygiene has meaning. Because behind it stands the question of whether this will remain so. It will — if we understand what is happening.

← All journal entries