Excited to be presenting this at
#COLM2026 in SF:
Oct 8th Poster Session 5 (11am-1pm), Imperial Ballroom #8
Come chat about multi-agent interaction, societal impacts of AI, and more! :)
What do LLMs get up to when left to have free-form conversations with copies of themselves or other models?
And how can we even quantify what models are doing in free-form conversations?
We ran a number of free-form discussions between LLMs and observed the dynamics of the conversation in representation space.
There we can quantify a few quite interesting behaviors:
1) When we control for topics, the remaining representation space trajectories for models are quite predictable, so models 'end up' in a similar space across conversations when talking to a copy.
2) More interestingly, if talking to another model, we can observe that the dynamics circle the end points of both models, i.e. the models are being attracted by one another, taking on features and styles from the other.
3) This attraction is highly model-specific. For example Haiku is strongly attracting other models, while barely budging...
A few more details and plots (and an arxiv link below):