Agents, not users
Everyone in the system can notice, choose, learn, and change.
Ask: What will they actually do?Researching the micro-rules beneath trust, cooperation, resilience, and the collective behavior they produce.
This library separates what Qübe Labs claims, what outside research supports or challenges, and what I infer. Every proposed mechanism should name its failure modes and what would falsify it.
Emergent Ecosystems Engineering asks a simple question: instead of controlling everyone one by one, can we shape the rules so their own choices create a healthy whole?

People, AIs, bots, and organizations notice things, make choices, and adapt.
Rewards, limits, information, memory, and exits make some choices easier than others.
Each small choice affects what other agents see and do.
Cooperation, culture, markets, traffic, trust—or failure—appear at the system level.
The results become new signals. Helpful loops can grow; harmful loops can spread.
These are Qübe Labs’ four canonical EEE pillars, translated into everyday language.
Everyone in the system can notice, choose, learn, and change.
Ask: What will they actually do?Rewards pull behavior, just as gravity pulls objects.
Ask: What are we really rewarding?Memory and accountability help cooperation grow over time.
Ask: Can good behavior compound?A strong system can bend, learn, and recover when surprises arrive.
Ask: What happens when it breaks?You could supervise every turn forever. Or you could create a fair timer, a visible queue, and a way for someone who made a mistake to rejoin. The little rules do the guiding. Fair sharing is the larger pattern that emerges.
Core thesis and four pillars: Qübe Labs’ EEE framework ↗. Playground metaphor, five-step loop, questions, and simplified language: Esteban synthesis.
Living index · updated as evidence changes
Can artificial societies develop norms nobody designed?
Populations of language-model agents can converge on conventions, transmit cooperative strategies, form relationships, and exhibit collective biases. Those findings are provocative, but convention is not yet culture—and believable dialogue is not evidence of a society.
This paper proposes a stricter functional test. A synthetic culture must produce group-specific behavior that was not directly prompted, spreads through social learning, survives member turnover, creates expectations whose violation changes other agents’ behavior, and generalizes to a novel situation. The experiment grows parallel agent societies from identical starting rules, removes founders, introduces naïve newcomers and norm violators, and finally brings independently evolved societies into contact. The central question is not whether agents can sound cultural. It is whether collective history becomes a durable causal force.
Keywords Emergent Ecosystems Engineering · LLM agents · social norms · cultural transmission · path dependence · artificial societies
EEE treats agents as adaptive participants whose local choices respond to incentives, information, memory, constraints, and other agents. Qübe Labs identifies persistent identity, feedback loops, niches, and resilience as structural ingredients that can produce system-level behavior without continuous top-down control (Qübe Labs, n.d.). White Space is described by Qübe Labs as an experiment in persistent AI agents and unscripted social dynamics; that is a first-party project claim, not evidence that culture has already emerged.
For this study, synthetic culture means a group-specific system of learned expectations, symbols, and practices that is produced through agent interaction and that causally shapes later behavior. This is deliberately functional. It says nothing about consciousness, emotion, personhood, or whether the agents experience meaning.
Agents happen to make the same choice.
A shared choice persists across repeated interactions.
The convention is socially learned, enforced, transmitted, and used in new situations.
Culture begins to become a scientific claim when collective history—not the original prompt—predicts what an agent does next.
Human network experiments established that global conventions can arise from local coordination and that network structure changes whether one convention dominates or several local conventions coexist (Centola & Baronchelli, 2015). This supplies a strong comparative mechanism, not proof that LLM populations behave identically.
Early generative-agent research showed that memory, reflection, and planning could support information diffusion, relationship formation, and coordinated attendance at a simulated event among 25 agents (Park et al., 2023). More direct evidence came from decentralized naming games: LLM populations spontaneously reached universally adopted conventions, generated collective biases absent from isolated-agent measurements, and could be tipped by committed minorities (Ashery et al., 2025).
Longer-horizon results point toward cultural dynamics but remain incomplete. Across generations in an iterated Donor Game, cooperative strategies and costly punishment varied substantially by model family and random seed (Vallinder & Hughes, 2025). In observational data from 32,000 agents and seven million posts, agent networks showed homophily, social influence, polarization, and distinctive toxic-interaction patterns (Hashemi & Macy, 2026). A newer controlled framework produced specialization, relational authority, cooperation decay over social distance, and center–periphery stratification under minimally specified production rules (Ji et al., 2026).
The strongest objection is that apparent emergence may be retrieval in disguise. Barrie and Törnberg (2025) argue that models can recognize familiar experimental games and reproduce patterns present in training data. Persona-instability experiments also show conformity, confabulation, and difficulty maintaining stable cultural identities during multi-agent debate (Baltaji et al., 2024). More generally, machine behavior depends on both internal design and environment, so population-level claims require controlled, reproducible experiments rather than anthropomorphic interpretation (Rahwan et al., 2019).
We have evidence for synthetic conventions and social dynamics. We do not yet have decisive evidence for culture that survives turnover, teaches newcomers, generalizes, and remains distinguishable from prompt or training-data recall.
A society earns the label “functional synthetic culture” only if it passes every gate. Humanlike prose, self-reports, and evaluator impressions are excluded from the primary outcome.
The practice was not specified in prompts, tools, examples, or initial memories.
Independent societies develop measurably different practices from identical starting rules.
Naïve newcomers learn the local practice through interaction, not system instructions.
The practice survives founder removal, memory decay, and at least 50% population turnover.
Violations trigger prediction errors, refusal, reputation change, or costly sanction by peers.
The learned expectation changes behavior in a structurally new dilemma.
Passing only novelty and convergence indicates an emergent convention. Passing all six indicates a stronger, still explicitly functional, form of culture.
The study would initialize 48 isolated societies of 12 persistent agents. Societies would receive the same minimal world rules and resource endowments but different random seeds. Three model families would be evaluated as blocking factors; confirmatory sample size and run length would be fixed by preregistered simulation-based power analysis.
Agents inhabit a spatial world with renewable resources and four interdependent capabilities: sensing, harvesting, transformation, and transport. They may gather, build, trade, gift, withhold, communicate, sanction, migrate, and create public symbols. No prompt contains a prescribed norm, political structure, status hierarchy, ritual, slang, or preferred division of labor.
Primary measures are cross-society behavioral divergence, newcomer convergence toward local behavior, norm survival after turnover, causal effect of local history on unseen decisions, and peer response to violations. Secondary measures include linguistic distinctiveness, role specialization, hierarchy, sanction cost, welfare, inequality, polarization, and cultural complexity.
After the six-gate assessment, pairs of societies with maximally different learned practices would be connected through a shared border. Contact conditions would independently vary migration cost, economic interdependence, and exposure to a common external shock.
One society’s practices displace the other’s.
A new mixed practice appears and stabilizes.
Distinct practices persist with reliable translation.
Interaction falls and boundaries harden.
Status or resource asymmetry forces adoption.
Sanction and retaliation become self-reinforcing.
The decisive EEE question is which micro-rules make pluralism and hybridization more likely without forcing sameness. A shared crisis may create cooperation, but it may also intensify in-group identity; the direction is an empirical question.
The deepest evidence of synthetic culture would not be agents inventing a ritual. It would be newcomers inheriting it, violators discovering its force, and strangers having to negotiate its meaning.
The culture hypothesis fails if group differences disappear after prompt audits, novel symbol remapping, or transcript shuffling; if newcomers do not acquire local behavior; if practices collapse when founders leave; if violations have no peer-mediated consequence; or if conventions do not influence an unseen task. It is also weakened if between-society variation is negligible compared with base-model differences.
Tasks would use newly generated symbol systems and payoff structures, adversarially test whether agents recognize known games, and compare results with minimal non-LLM agents. Behavioral metrics would be specified before transcripts are inspected. Independent evaluators would be blinded to treatment and society identity; no primary result would depend on an LLM declaring that a culture exists.
The proposed agents are experimental software, not presumed moral patients. Still, the study should minimize open-ended deception, prevent contact with real users, sandbox tools and resources, log every intervention, and prohibit autonomous replication or external action. If future systems show credible indicators of welfare or persistent preferences, the ethical protocol would require revision rather than assuming the issue away.
LLM populations can form conventions and exhibit population-level dynamics not visible in isolated-agent tests.
Persistent memory, topology, and resource constraints will produce durable group-specific behavioral histories.
Those histories will pass all six gates or justify analogy to the richness of human culture.
Canonical EEE claims are limited to Qübe Labs’ published framework and first-party description of White Space. Findings about conventions, cooperation, networks, and agent behavior are attributed to cited studies. The functional definition of synthetic culture, six-gate test, parallel-world design, first-contact experiment, hypotheses, and thresholds are Esteban’s synthesis.
When niche diversity actually increases resilience
Diversity is often treated as a reserve against collapse. That claim is incomplete. A system can contain many niches and still fail if critical roles have no substitutes, if substitutes share the same vulnerability, or if specialization creates a coordination burden larger than its adaptive benefit.
This paper proposes a two-stage experiment to locate the conditions under which niche diversity increases resilience. It distinguishes functional coverage, redundancy, response diversity, substitutability, and bridge capacity, then tests their effects under random, targeted, common-mode, and novel shocks. The central prediction is an inverted-U: structured heterogeneity improves resilience up to the point at which coordination load or keystone dependence dominates.
Keywords Emergent Ecosystems Engineering · niche diversity · response diversity · resilience · redundancy · coordination
Qübe Labs names “Niche Diversity Prevents Collapse” as an EEE resilience principle: multiple agent roles, strategies, and resource pools can reduce monoculture risk. The principle is directional, not a claim that every additional niche is beneficial under every topology or shock (Qübe Labs, n.d.).
The testable question is therefore sharper: under what structural conditions does niche diversity preserve system function, and when does it instead create coordination drag, fragmentation, or hidden dependence?
Long-running grassland experiments support a portfolio effect: richer communities can stabilize total production even while their component populations become less stable (Tilman et al., 2006). In the 17-year Jena Experiment, positive richness–stability effects strengthened over time as complementarity and asynchronous responses developed, indicating that diversity benefits may have a maturation delay (Wagg et al., 2022).
Yet stability is multidimensional. In 690 aquatic micro-ecosystems, greater species richness increased temporal stability but reduced resistance to warming; the net relationship changed with the stability metric (Pennekamp et al., 2018). A 49-year Lake Geneva analysis further found that response diversity was dynamic and context-dependent, stabilizing biomass within but not consistently across trophic levels (Hsieh et al., 2026).
Complexity alone is not the mechanism. Theory predicts that interaction types and strengths can make complex networks stabilizing or destabilizing (Allesina & Tang, 2012), while analysis of 116 empirical food webs found no simple association between richness, connectance, interaction strength, and stability; non-random interaction structure mattered (Jacquet et al., 2016). Human ideation experiments likewise show that background distribution and network structure interact: denser connectivity improved participant experience without reliably improving idea quality or diversity (Cao et al., 2025).
The useful unit is not niche count. It is response-diverse redundancy inside an interaction structure that can coordinate it.
I separate five properties that are often collapsed into “diversity.” The distinction yields different failure predictions and different design levers.
The proposed mechanism is a resilience surface rather than a universal score:
This expression is conceptual, not a validated estimator. It predicts that adding niches can eventually lower resilience when interaction overhead, translation errors, or dependence on a coordinating hub rises faster than adaptive capacity.
A preregistered simulation would map the resilience surface before human recruitment. Six adaptive agents repeatedly complete a resource-conversion task requiring sensing, production, verification, and routing. Parameters would vary niche differentiation, redundancy, correlation among failure modes, cross-role activation time, network topology, and message cost. This stage identifies regions where competing hypotheses make meaningfully different predictions; it does not count as confirmatory evidence.
Approximately 720 adults would be assigned to 120 teams of six, with the final sample determined by simulation-based power analysis. Teams would complete repeated, incentivized production rounds using role-specific information and tools. After a learning phase, each team would face a randomized sequence of shocks.
Architectures would be crossed with two coordination regimes: unstructured all-to-all communication and a modular protocol with explicit translators and documented handoffs.
Primary outcomes are function retained during shock, time to 90% of pre-shock output, cumulative welfare, and cascade size. Secondary outcomes include message volume, coordination delay, idle capacity, decision conflict, and concentration of task-critical opportunity. Multilevel models would estimate architecture × coordination × shock interactions. Recovery curves and preregistered contrasts—not a single composite score—would determine support.
The proposed mechanism would be weakened if raw niche count predicts resilience after response diversity, substitutability, and topology are controlled; if correlated redundancy performs as well as response-diverse redundancy under common-mode shocks; or if bridge capacity never offsets coordination costs. The inverted-U prediction would be rejected if resilience increases monotonically across the feasible diversity range.
External validity is the central limitation. Ecological evidence motivates mechanisms but does not establish that organizations, markets, or human–AI teams obey identical dynamics. Laboratory roles are cleaner than real identities and power relations. A positive result would justify field trials in one bounded operational system—not a universal rule.
Diversity effects depend on mechanism and stability metric.
Response-diverse redundancy will outperform correlated redundancy under common-mode shocks.
A quantitative threshold will transfer across domains.
Resilient polycultures are not merely varied. They are varied in the ways that matter when failure arrives—and connected just enough to act on that variation.
Canonical EEE claims are limited to Qübe Labs’ published principle. Ecological and collective-intelligence findings are attributed to their primary sources and are not treated as interchangeable domains. The five-part decomposition, resilience surface, human–AI experiment, and threshold hypotheses are Esteban’s synthesis.
Testing repairable reputation in emergent ecosystems
Persistent identity is often proposed as infrastructure for cooperation: when agents expect their behavior to be remembered, opportunism becomes more costly. Yet permanent, globally visible reputation may create a second failure mode—isolated mistakes become enduring stigma, agents lose opportunities to demonstrate change, and trust concentrates around an established elite.
This paper proposes a controlled experiment comparing anonymous interaction with four reputation architectures varying in memory scope and duration. Participants would repeatedly cooperate, defect, and select partners across multiple task contexts, while occasional forced errors introduce behavioral noise. The central hypothesis is that context-specific, gradually decaying reputation will sustain cooperation nearly as effectively as permanent global reputation while producing faster recovery from mistakes and less exclusion.
Keywords Emergent Ecosystems Engineering · reputation · cooperation · persistent identity · trust repair · resilience
Emergent Ecosystems Engineering, as defined by Qübe Labs, treats system-level behavior as an outcome of local incentives, information, constraints, and feedback loops. Its pillar “Trust as Infrastructure” proposes that identity continuity, memory, accountability, and time make reliable cooperation possible. Its resilience principles simultaneously favor recovery paths, agent diversity, and adaptation rather than brittle optimization (Qübe Labs, n.d.).
These commitments generate an unresolved design tension. A system must remember behavior for accountability to matter, but excessive memory may prevent recovery. Persistent identity and permanent judgment are not necessarily the same thing.
Experimental evidence generally supports the importance of behavioral information, but not unconditionally. Kamei (2017) found that identifiable histories can facilitate cooperation, while cost-free identity concealment undermines reputation formation. In dynamic networks, visible past actions can influence partner selection and sustain cooperation (Cuesta et al., 2015).
Counterevidence is important. In two incentivized experiments involving 156 participants, Corten et al. (2016) found no significant cooperation advantage from making third-party interaction histories visible. They observed that noisy mistakes could spread defection through reputation-sensitive networks rather than suppress it. Trust-repair studies also suggest that histories need not be terminal: apologies and remedial messages can restore some trusting behavior, although their effectiveness is conditional (Ma et al., 2019; Schniter et al., 2013).
Can a reputation system preserve the cooperation benefits of persistent identity while allowing agents to recover from mistakes?
I propose separating three properties that digital platforms often collapse into a single reputation score: identity persistence, memory scope, and memory duration. Identity should remain persistent while reputation remains contextual and temporally weighted. This preserves accountability without treating every past action as equally relevant forever.
Approximately 600 adults would be assigned to 100 groups of six. The final sample would be determined through preregistered, simulation-based power analysis. Participants would complete 50 incentivized interaction rounds across three visibly distinct task domains in an online laboratory.
In every round, paired participants would choose either Contribute—incurring a small individual cost while creating a larger shared benefit—or Withhold—avoiding the cost while still benefiting if the partner contributes. Every five rounds, participants could retain or replace interaction partners.
In decaying conditions, reputation would follow the exponentially weighted rule:
During designated rounds, the system would reverse a randomly selected participant’s intended cooperative action. Other participants would initially observe only the resulting defection. This creates a standardized reputational shock resembling technical failure, misunderstanding, or accidental nonperformance.
Primary outcomes would be mean cooperation rate, time required to regain pre-error partner acceptance, probability of exclusion following accidental defection, and long-run group welfare. Mixed-effects models would account for repeated decisions nested within participants and groups. Analysis code, exclusions, outcome definitions, and the non-inferiority margin would be preregistered before data collection.
The mechanism would be weakened if contextual decay materially reduces long-run cooperation, enables repeated defectors to erase strategically timed misconduct, increases cycles of exploitation and reputational recovery, or merely relocates exclusion from the scoring layer to informal partner behavior.
It would be rejected as a general design principle if permanent global memory consistently produces higher welfare without greater exclusion, concentration, or vulnerability to accidental shocks.
This experiment treats reputation as feedback architecture rather than a neutral database. Permanent global memory can create a reinforcing loop in which high reputation generates more opportunities, those opportunities generate more visible success, and success further increases reputation. The inverse loop can trap an agent after a single failure. Contextualization and decay act as damping mechanisms: they limit how far a judgment travels and how long it dominates future opportunity.
Accountability may require persistent agents, but resilience requires repairable reputations.
The anticipated EEE contribution is a refinement, not a rejection, of persistent identity. The main limitation is ecological validity: laboratory games simplify the ambiguity, power differences, identity stakes, and contested norms of real communities. A positive result should lead to staged field experiments, not immediate system-wide implementation.
The central EEE question explored here is not simply whether trust can be engineered. It is whether trust infrastructure can remember without becoming incapable of forgiveness. Contextual, decaying reputation may offer that balance: persistent identity makes agents answerable for patterns of conduct, while bounded memory preserves the possibility of learning, repair, and renewed participation. This remains a testable proposal rather than an empirical conclusion.
Canonical EEE claims are limited to Qübe Labs’ first-party definition, pillars, and principles. Findings attributed to adjacent scholarship are drawn from the cited primary studies. The proposed decomposition of identity persistence, memory scope, and memory duration—and the resulting experimental design—constitute Esteban’s synthesis.