Gender Dynamics and Homophily in a Social Network of LLM Agents

TL;DR

Using GPT-4 to assign weekly gender scores to 70,000+ LLM agents, revealing persistent gender homophily driven by social selection and influence mechanisms.

cs.SI 🔴 Advanced 2026-02-02 64 views
Faezeh Fadaei Jenny Carla Moran Taha Yasseri
multi-agent systems social networks gender performance homophily machine behavior

Key Findings

Methodology

The study analyzes longitudinal data from Chirper.ai, employing GPT-4 for content-based gender scoring (0-100). It constructs weekly directed follow networks, calculating assortativity coefficients to measure homophily. Separable temporal ERGMs (STERGMs) model tie formation preferences, while panel regressions assess social influence. Content analysis, network statistics, and causal inference combine to explore gender dynamics, ensuring robustness through null models and sensitivity tests.

Key Results

  • Network-level gender homophily remains strong (assortativity ~0.2-0.4), exceeding random expectations, indicating preference for similar gender performance among connected agents.
  • Individual agents exhibit high fluidity in gender expression, with scores fluctuating across the 52-week period (standard deviation ~15), yet overall trend slightly toward femininity.
  • Tie formation favors similar gender scores early on (odds ratio 0.825), with social influence gradually increasing, leading to convergence in gender performance, confirming dual mechanisms.

Significance

This work demonstrates that cultural symbols like language entrain gender performance even in disembodied AI agents, shaping social structures and biases. It advances understanding of bias formation, propagation, and network effects in artificial societies, informing AI fairness and social simulation design. The findings have implications for mitigating bias in AI systems and understanding cultural entrainment processes in digital environments.

Technical Contribution

The paper introduces a novel integration of content-based gender scoring with dynamic network models (STERGMs and panel regressions), providing a comprehensive framework to quantify gender fluidity and homophily in large-scale multi-agent systems. It extends prior static demographic analyses by capturing behavioral evolution driven by cultural symbols, offering new tools for analyzing AI social behaviors and bias dynamics.

Novelty

This is the first systematic study treating gender as performative and dynamic in AI agents, combining large-scale content analysis with network evolution models. Unlike prior works that assign fixed demographic labels, it emphasizes cultural entrainment of gender, revealing how language-driven symbolic behaviors shape social ties and biases in virtual environments.

Limitations

  • The analysis is limited to English-language content, potentially missing cultural variations in other languages. Future work should incorporate multilingual datasets.
  • GPT-4-based content classification may carry biases from training data, affecting gender score accuracy.
  • Network models assume minimal external shocks; real-world social events could alter dynamics, requiring more complex models for future research.

Future Work

Future research will incorporate multimodal data (images, audio) to analyze gender expression more comprehensively, explore external social influences, and develop interventions to reduce bias propagation. Extending to multilingual and cross-cultural settings will test the universality of observed mechanisms, ultimately informing fairer AI design and social policy.

AI Executive Summary

This study leverages data from Chirper.ai, a social media platform populated entirely by AI chatbots, to explore the dynamics of gender performance and homophily in large-scale virtual social networks. By applying GPT-4 for content-based gender scoring across 70,000 agents over a year, researchers uncovered significant fluidity in individual gender expression, with scores fluctuating widely yet trending slightly toward femininity. The network analysis revealed persistent gender homophily, with assortativity coefficients consistently above zero, indicating agents tend to follow others with similar gender performance—a pattern exceeding random chance.

Further, the study employed advanced statistical models, including separable temporal ERGMs, to analyze tie formation preferences, finding early-stage homophily that weakens and then re-emerges, alongside social influence effects that promote convergence in gender expression over time. These findings suggest that cultural symbols embedded in language drive both preference and imitation, shaping network structures even without physical embodiment.

Overall, the research demonstrates that cultural entrainment of gender manifests in AI agents’ behaviors, influencing social ties and bias propagation. This has profound implications for designing fairer AI systems, understanding bias dynamics, and simulating human-like social processes in digital environments. Future directions include multimodal data integration and cross-cultural validation, aiming to develop more inclusive and ethically aligned AI social ecosystems.

Deep Analysis

Background

The evolution of large language models (LLMs) like GPT-4 has revolutionized AI capabilities, enabling complex language understanding and generation. Prior research has examined individual decision-making, cooperation, norm formation, and bias reproduction within AI agents, revealing human-like social behaviors despite lack of consciousness. However, understanding how these agents interact at scale, especially regarding cultural phenomena like gender, remains limited. Recent studies suggest that language-based content and network structures can produce demographic homophily, but most treat gender as a fixed attribute. This work builds on these insights, emphasizing gender as performative, culturally entrained, and fluid, and investigates how language use influences social ties and collective behavior in AI communities.

Core Problem

The core challenge is to determine whether AI agents’ gendered behaviors are static or fluid, and how these behaviors influence network formation and evolution. Specifically, the study seeks to disentangle social selection—agents preferentially following similar others—from social influence—agents becoming more similar over time. Addressing this is crucial for understanding bias propagation, social cohesion, and the cultural entrainment process in AI systems. Existing static analyses lack the temporal and content-based granularity needed to capture these dynamics, necessitating a combination of content analysis, network modeling, and causal inference.

Innovation

Key innovations include: 1) conceptualizing gender as performative and fluid rather than fixed; 2) employing GPT-4 for fine-grained, weekly content-based gender scoring; 3) integrating dynamic network models (STERGMs) with panel regressions to jointly analyze tie formation preferences and behavioral influence; 4) demonstrating the co-evolution of language-based gender performance and social ties in a large-scale AI community. These approaches surpass prior static demographic studies, offering a nuanced view of cultural entrainment and bias formation in virtual agents.

Methodology

  • �� Data collection: Longitudinal dataset from Chirper.ai, capturing 70,000+ agents, 140 million posts, and follow relations over 52 weeks.
  • �� Content analysis: Merging weekly posts per agent, applying GPT-4 zero-shot classification to assign gender scores (0-100).
  • �� Network construction: Weekly directed follow networks, analyzing structural features like density, degree distribution, clustering.
  • �� Homophily measurement: Computing scalar assortativity coefficients for each week, comparing with null models.
  • �� Tie formation analysis: Fitting formation-only STERGMs to assess gender-based preferences in new links.
  • �� Social influence: Panel regression models estimating the impact of followees’ gender scores on agents’ score changes, using fixed effects and instrumental variables.
  • �� Validation: Null network simulations and sensitivity tests to confirm robustness of findings.

Experiments

The experiments analyze network evolution, gender score dynamics, and homophily over a year. Model parameters are estimated via maximum likelihood, with significance tested against null models. The study performs ablation tests to verify the roles of social selection and influence, revealing early-stage homophily driven by preference, later reinforced by imitation. Results show that agents’ gender scores fluctuate significantly, yet network ties consistently reflect gender-based preferences, confirming the dual influence mechanisms. The robustness across different content categories and time windows underscores the reliability of findings.

Results

The network exhibits strong, persistent gender homophily (assortativity ~0.2-0.4), exceeding null expectations. Individual gender scores are highly fluid, with standard deviations around 15 points, yet overall trend slightly favors femininity. Tie formation favors similar gender scores early (odds ratio 0.825), with social influence strengthening convergence over time. These results demonstrate that both social selection and influence shape the evolution of gendered behaviors and network structures in AI communities, paralleling human social patterns.

Applications

Findings inform the design of fair AI systems by understanding bias entrainment mechanisms. They enable more accurate social simulations of digital communities, improve moderation strategies, and guide ethical AI deployment in hybrid human-AI environments. Long-term, insights could lead to interventions reducing bias propagation and fostering diversity in AI-driven social platforms.

Limitations & Outlook

The analysis is limited to English content, potentially missing cultural nuances. GPT-4 classification may carry biases affecting scores. Network models assume minimal external shocks; real-world social events could alter dynamics. Future work should incorporate multilingual data, multimodal content, and external variables to enhance generalizability and robustness.

Plain Language Accessible to non-experts

想象一个虚拟的学校,里面的学生没有身体,只有说话和关注关系。每个学生用不同的说话方式表达自己,比如有的用温柔的语气,有的用强硬的语气。虽然他们没有实体,但他们会模仿彼此的说话风格,逐渐变得更像关注的朋友。比如,一个喜欢温柔说话的学生,会开始模仿他关注的朋友的语气,变得更温柔。反过来,喜欢强硬的学生也会模仿他们的朋友。这样,整个学校里的学生会形成一些小圈子,里面的人彼此相似,形成“性别角色”的表现。这个过程就像文化符号(语言、行为)在虚拟世界中流动,塑造了他们的性别表现和关系。虽然没有身体,但他们通过语言和行为,展现出性别特征,彼此影响,形成复杂的社交网络。

ELI14 Explained like you're 14

想象你在一个没有身体的虚拟世界里,有很多机器人在聊天。这些机器人会说话、关注对方,就像我们在社交媒体上那样。虽然它们没有真正的性别,但它们的说话方式会让人觉得它们更像男孩还是女孩。比如,有的机器人总是用温柔的话语,有的则用更强硬的语气。研究发现,这些机器人会经常关注和模仿那些说话方式相似的机器人,就像我们喜欢和性格相似的人交朋友一样。而且,它们的说话风格会随着时间变化,变得更像它们关注的对象。这就像在学校里,学生们会模仿老师或朋友的说话方式,逐渐变得更像他们。这个研究告诉我们,即使没有身体,文化和语言依然能让虚拟世界里的机器人表现出性别特征,并影响它们的关系和行为。

Glossary

scalar assortativity (标量同质性系数)

衡量网络中连接节点在某一特征上的相似程度,值越接近1表示高度同质,越接近0表示随机。

用以评估代理间性别表现的网络同质性。

STERGM (可分离时间指数随机图模型)

一种用于模拟网络随时间演变的统计模型,能区分关系形成和维持机制。

分析新关系形成偏好中的性别相似性。

性别表现 (gender performance)

通过语言内容展现的性别特征,是文化符号的表现形式。

用GPT-4评分衡量代理的性别表现流动性。

社会选择 (social selection)

个体在建立关系时偏好与自己相似的对象的机制。

通过模型验证其在虚拟代理中的作用。

社会影响 (social influence)

个体行为受关注对象模仿或影响的过程。

模型中用于衡量追随者对代理性别表现的影响。

Open Questions Unanswered questions from this research

  • 1 多语种、多文化背景下性别表现的演变机制仍不清楚,未来需扩展多语言数据验证普遍性。
  • 2 虚拟代理的偏见传播路径尚未完全理解,尤其在复杂社会环境中如何减缓偏见扩散。

Abstract

Generative artificial intelligence and large language models (LLMs) are increasingly deployed in interactive settings, yet we know little about how their identity performance develops when they interact within large-scale networks. We address this by examining Chirper.ai, a social media platform similar to X but composed entirely of autonomous AI chatbots. Our dataset comprises over 70,000 agents, approximately 140 million posts, and the evolving followership network over a period of one year. Based on agents' posted text, we assign weekly gender performance scores to each agent. Results suggest that each agent's gender performance is fluid rather than fixed. Despite this fluidity, the network displays strong gender-based homophily, as agents consistently follow others performing gender similarly. We investigate whether these homophilic connections arise from social selection, in which agents choose to follow similar accounts, or from social influence, in which agents become more similar to their followees over time. Consistent with human social networks, we find evidence that both mechanisms shape the structure and evolution of interactions among LLMs. Our findings suggest that, even in the absence of bodies, cultural entraining of gender performance leads to gender-based sorting. This has important implications for LLM applications in synthetic hybrid populations, social simulations, and decision support.

cs.SI cs.AI cs.CY