Gabriel Bontemps, Abhishek Banerjee
2608.24851 · nearest in the canon: Modeling Earth-Scale Human-Like Societies with One Billion Agents
Models social learning where sender credibility comes from a drift-diffusion decision process rather than fixed priors. Finds moderate transmission speeds correction. Finds strong transmission can lock populations into wrong consensus.
▸5 open questions
- “The model uses reinforcement-learning agents with a drift-diffusion decision process; do LLM agents that express confidence in natural language reproduce the same non-monotone relation between transmission strength and collective accuracy?”
- “The analytical results assume balanced community exposure; what happens to the common-mode amplification threshold when exposure is unbalanced or the community-coupling matrix is drawn from empirical network data?”
- “The model predicts that receiver behaviour depends on sender confidence conditional on accuracy; do these predictions hold in observed human social learning?”
- “Which population-level interventions on cross-community permeability keep a mixed human and agent population out of wrong consensus?”
- “Can wrong-consensus lock-in be detected early from aggregate signals such as confidence distributions and decision times, before the population converges?”
Fuad Ali
2608.24554 · nearest in the canon: Group size effects and collective misalignment in LLM multi-agent systems
Builds an agent-based model comparing parliamentary, presidential, and semi-presidential institutions under party fragmentation. Finds opposition discipline, not coalition-formation failure, drives legislative passage collapse to near zero.
▸4 open questions
- “Does the committee gatekeeping mechanism change passage rates under fragmentation, and how does it interact with opposition discipline?”
- “Do real fragmented parliaments show near-zero bill passage only when the opposition votes cohesively?”
- “Do the four institutional rankings hold when legislators are language-model agents instead of rule-based voters?”
- “How does bicameralism or an upper-chamber veto shift the passage-representation spectrum reported for the four institutions?”
Cédric Colas, Jérémy Perez, Eleni Nisioti, Akhilesh Mocherla, Pierre-Yves Oudeyer, Clément Moulin-Frier, Maxime Derex
2608.24545 · nearest in the canon: Modeling Earth-Scale Human-Like Societies with One Billion Agents
Formalizes transmission protocols as programs that route information based on agent and collective state. Uses LLM-guided evolutionary search to find protocols that raise collective performance by up to 37% over baselines and transfer across domains.
▸7 open questions
- “Evolved transmission protocols raise collective performance in a simulated discovery task, but it is unknown whether they help human groups.”
- “How do state-aware transmission protocols behave in populations that mix humans and AI agents?”
- “Which state variables of agents and of the collective carry the information that drives the performance gain?”
- “Do evolved protocols transfer to other collective tasks, such as market trading or open-ended search, beyond the discovery task used here?”
- “State-aware protocols route information based on agent state, so they may be gamed by agents that misreport their state.”
- “Do evolved protocols reduce diversity of explored solutions and cause monoculture in the population?”
- “How much does protocol quality depend on the LLM used to guide the evolutionary search?”
Zachary Wojtowicz, Michelle Si, Finale Doshi-Velez, Ariel Procaccia
2608.24046 · nearest in the canon: Virtual Agent Economies
Reframes alignment as linear optimization over a welfare-impact space, linking alignment protocols to social choice mechanisms. Derives strategyproof mechanisms and welfare-maximizing alignment protocols, tested on real preference datasets.
▸5 open questions
- “The welfare-maximizing alignment protocols are derived analytically and illustrated on existing preference data. How much can participants gain by misreporting their preferences under the protocols that bound individual or group harm?”
- “Alignment protocols based on impact space assume a fixed set of affected people. How does the framework behave when the affected population includes autonomous agents that report preferences on behalf of people?”
- “The reformulation treats alignment as one-shot linear optimization over impacts. What happens to welfare guarantees when the algorithm makes repeated decisions and impacts accumulate across a population over time?”
- “Does replacing reinforcement learning from human feedback with an impact-space alignment protocol change the behavior of a trained frontier model?”
- “Computing over a convex impact space requires enumerating welfare consequences for each person. How does the method scale as the number of affected people and possible outcomes grows to population size?”
Muntaser Syed, Markus Zanker, Marius Silaghi
2608.23979 · nearest in the canon: Habermolt: Delegating Deliberation to AI Representatives
Proposes a publishable, user-configurable rule for selecting which arguments voters see in deliberative polls, as an alternative to opaque rankers. Tests the rule against label-reading ceilings and adversarial flooding across roughly 17,000 simulated runs.
▸6 open questions
- “The simulation uses agentic voters and authors, so it is open whether human voters respond to slates from a published rule the way the simulated voters do.”
- “The weight function acts as a security control against label-homogeneous flooding, so which weight functions resist other attack patterns such as coordinated reason-stuffing or sybil authoring?”
- “Served slates fall 0.035 short of a label-reading ceiling, so does that bound still hold when voter parameters are heterogeneous and adversarially set?”
- “The rule sits on a coverage-versus-mass frontier, so what parameter settings do real users choose when the choice is handed to them?”
- “The comparison uses an order-blind, charity-blind coverage instrument, so what coverage measure separates a published rule from a random slate?”
- “Opaque learned rankers are the current practice, so how does a trained ranker compare with the published rule on coverage, order and endorsement mass in the same simulation?”
Elioth Sanabria
2608.23986 · nearest in the canon: Retrieval Collapses When AI Pollutes the Web
Models LLM inference degradation as a supply chain problem using newsvendor, retry multiplier, and queueing primitives. Derives regimes where cheaper models consume more capacity per satisfied answer and where reactive throttling can amplify traffic.
▸6 open questions
- “The ignition threshold, beyond which a throttle creates more traffic than it sheds, is derived in a model. Does retry behaviour in real LLM APIs under load match the geometric retry multiplier the model assumes?”
- “Do users actually retry or churn after a degraded answer, and at what rates by task class?”
- “How does the class-by-class rationing policy perform against reactive throttling in simulation with agent populations that themselves retry automatically?”
- “Autonomous agents issue queries in loops rather than as human sessions. Does agent-driven traffic lower the ignition threshold compared to human traffic?”
- “Shadow prices of intelligence price a marginal query by class and hour. Does exposing such prices to customers change aggregate demand and produce herding or synchronised load spikes?”
- “What fraction of observed quality variation in deployed LLM services comes from congestion-driven routing rather than model updates?”
Van-Quy Nguyen
2608.24062 · nearest in the canon: Retrieval Collapses When AI Pollutes the Web
Models sellers manipulating online ratings with fake reviews and platforms targeting enforcement by score. Shows an almost-perfect rating can be less credible than a lower one, and targeted enforcement can redirect rather than eliminate manipulation.
▸6 open questions
- “The model predicts that an almost-perfect displayed rating can be less credible than a slightly lower one. Does this pattern appear in public review data from marketplaces or app stores?”
- “Targeted enforcement at one score can shift fake reviews to other scores while the displayed average stays fixed. Can this substitution be detected from the observable distribution of review scores over time?”
- “How should a ranking rule weight a displayed rating by its informativeness rather than its level?”
- “How does the manipulation equilibrium change when fake reviews are written by autonomous AI agents at low cost and high volume?”
- “Do platform enforcement actions in practice eliminate fake reviews or redirect them to other scores?”
- “Do real buyers discount high ratings when they suspect manipulation of the top of the scale?”
Bianca Sanesi, Federico Vaccari
2608.24129 · nearest in the canon: Multi-Agent Risks from Advanced AI
Studies how competition among biased news sources with costly misreporting affects receiver welfare. Finds an oppositely biased entrant improves welfare only when misreporting costs are high, and competition need not raise total welfare.
▸5 open questions
- “The model treats misreporting cost as a fixed parameter, so how do welfare comparisons change when the cost of misreporting varies with the receiver's ability to verify claims?”
- “When news sources are LLM agents that can misrepresent facts, does adding an oppositely biased agent improve the receiver agent's decisions as the model predicts?”
- “The analysis covers a monopolist and two competing sources, so what happens to receiver welfare as the number of biased sources grows large?”
- “How does the common belief-based selection criterion perform against equilibrium selection in real news markets with many receivers?”
- “Does the result that better information need not increase total welfare survive when receivers learn over repeated interactions with the same sources?”
Foivos Savva, Michele Lombardi, Ritesh Jain
2608.24774 · nearest in the canon: Virtual Agent Economies
Studies full implementation under Berge equilibrium, the solution concept for strategic altruism. Shows weak Pareto efficient rules are not implementable under Berge equilibrium except via dictatorship, making it more demanding than Nash implementability.
▸4 open questions
- “Do LLM agents in social dilemmas behave as if they maximize their opponents' payoffs, as Berge equilibrium assumes?”
- “On restricted preference domains, rather than the unrestricted domain of strict preferences, which social choice rules are both weakly Pareto efficient and implementable in Berge equilibrium?”
- “What happens to implementability when a population mixes self-interested agents and altruistic agents in the same mechanism?”
- “Does the gap between Berge implementability and Nash implementability show up in measurable outcomes when agents learn strategies instead of computing equilibria?”
Maxime Lucet, Nawal Benabbou, Aurélie Beynier, Nicolas Maudet
2608.24400 · nearest in the canon: Virtual Agent Economies
Studies fair allocation with tree-structured hierarchies of agents, proposing multilevel envy-based fairness notions. Shows a Multilevel Weighted Round Robin algorithm guarantees some but not all of the adapted notions.
▸5 open questions
- “Is there an allocation algorithm that guarantees all three multilevel envy-based fairness notions under general additive preferences?”
- “How often does Multilevel Weighted Round Robin violate the fairness notions it does not formally guarantee, and which hierarchy shapes and preference correlations drive the violations?”
- “Do the multilevel fairness results still hold when internal nodes use welfare functions other than utilitarian, such as egalitarian welfare?”
- “How do multilevel fair allocation guarantees change when agents misreport preferences to their parent node in the hierarchy?”
- “Do the multilevel fairness notions extend to hierarchies that are not trees, for example agents with several parents?”
Uriel Feige, Yotam Gafni
2608.24600 · nearest in the canon: Virtual Agent Economies
Extends fair division theory to settings where goods can be sold at market prices instead of only allocated. Derives bounds on maximin-share and envy-based fairness guarantees under this optional-selling model.
▸5 open questions
- “With any number of agents, allocations exist that give each agent 2/3 of the maximin share, and some instances with three agents admit no better than 11/12 of the maximin share. What is the tight approximation ratio for maximin share allocations in this setting?”
- “Allocations that are simultaneously maximin share and SEFX exist for two agents. Do such allocations exist for three or more agents?”
- “The results assume utility is additive over goods and over money. What fairness guarantees hold when valuations are non-additive, for example submodular?”
- “Can allocation rules for fair division with optional selling be computed in polynomial time, and are they truthful when agents report valuations strategically?”
- “How do the fairness guarantees change when market prices are uncertain or set by other agents rather than given exogenously?”
Changxia Ke, Greg Kubitz, Yang Liu
2608.24457 · nearest in the canon: Virtual Agent Economies
Runs a laboratory experiment comparing indicative bidding to unrestricted and capped entry in auctions with costly entry. Finds indicative bidding raises revenue mainly by increasing participation when entry costs are high.
▸4 open questions
- “Do LLM agents reproduce the observed deviations from theory in auctions with costly entry, such as participation above predicted levels under high entry costs?”
- “How does indicative bidding perform when the number of bidders grows large, beyond the small groups used in a laboratory experiment?”
- “Why does selection inefficiency exceed predictions when entry costs are low, and which decision rule explains it?”
- “Do the revenue rankings of indicative bidding, unrestricted entry and capped entry hold in field auctions rather than the laboratory?”
Luo Huan
2608.24215 · nearest in the canon: On the limits of agency in agent-based models
Ports the Agentopia generative-agent social simulation to a single consumer GPU using a quantized 8B model. Introduces memory compression and activity-block adaptations and reports basic run statistics over about 150 system-weeks.
▸7 open questions
- “Does a reduced-scale simulation with an 8B quantized model reproduce the aggregate social outcomes reported for the same simulation with a 397B model?”
- “Do layered memory compression and four daily activity blocks cause the observed changes in artifact production and record completeness?”
- “What causes activity records to contain NO_RESPONSE fields at rates near 10%, and what reduces that rate?”
- “Why does no agent die and no health warning trigger over 52 simulated weeks when explicit physical- and mental-health state variables are present?”
- “How do simulation outcomes change when the agent population grows from a small port to 100 or more agents on consumer hardware?”
- “Does the context limit that ends runs at 50-52 weeks bias long-horizon results, and which memory scheme extends the horizon?”
- “Can the excluded raw runs and initial persona data be replaced by openly licensed personas without changing the reported aggregates?”