Scientists create an AI model with 8.3 bln personas

Scientists create an AI model with 8.3 bln personas
Фото: AI illustration

Researchers from Harvard University and the Massachusetts Institute of Technology (MIT) have presented MatrAIx, an infrastructure for testing digital products with the help of virtual users. At its core is the Persona 8B database, containing 8.3 bln records representing different combinations of human characteristics, El.kz reports citing a study published on arXiv.

How the virtual population works

Persona 8B contains 8.3 billion records generated using statistical data and multiple sources of information about people. The authors emphasize that the system is not intended to reconstruct specific individuals; rather, it models the diversity of human characteristics and behavior.

Each persona is described using 1,290 categorical parameters. These cover age, region, education, occupation, psychological characteristics, skills, habits, interests, risk attitudes, and patterns of interaction with technology.

Some profiles are generated synthetically based on statistical relationships between characteristics. Others are derived from anonymized information obtained from biographies, reviews, social research, and voluntary questionnaires.

What AI personas can do

MatrAIx enables virtual users to operate in four environments. They can complete surveys, interact with chatbots, visit websites, and use applications.

Researchers can define a target audience and a specific scenario, after which selected personas are assigned a task. The system records their actions, task completion time, outcomes, and responses to events occurring during the interaction.

This approach could potentially enable digital products to be tested before launch. For example, researchers can assess how different user groups respond to a price change, determine how convenient a particular application feature is, or examine whether a user continues interacting with a chatbot after an error occurs.

Why the project has attracted research interest

Traditional human-subject research requires time to recruit participants, organize surveys, and process results. With MatrAIx, parallel tests involving virtual users can produce initial results within several hours.

The project’s authors have created a library of 1,010 ready-made research tasks covering more than 25 fields, including commerce, software, finance, and healthcare. However, the inclusion of a task in the library does not necessarily mean that it has been experimentally conducted.

As part of the study, the team conducted 18,189 trials across eight tasks. The experiments employed three language models to control the virtual personas.

How realistically do the virtual people behave?

The researchers separately evaluated how consistently the agents adhered to the characteristics specified in their personas. In a controlled experiment involving 400 trials, the specified behavior was either expressed or appropriately suppressed in 91.5% of cases.

At the same time, the authors of the study caution that the technology has a significant limitation. A plausible AI response does not, in itself, demonstrate that a virtual user actually behaves in the same way as a real person.

Agent behavior is influenced by the language model being used. Therefore, simulation results must be interpreted in light of which model controlled the virtual personas and how the sample was constructed.

Can virtual simulations replace real-world surveys?

MatrAIx is primarily positioned as a tool for preliminary testing. It enables researchers to identify potential product issues quickly, compare different user groups, and repeat the same scenario after changes have been introduced.

At the same time, the researchers do not propose eliminating human participation altogether. The study emphasizes that human evaluation remains necessary to determine the extent to which simulation results correspond to real-world behavior.

El recommends