Method under test
Evolving AI Agent Collective
We are developing a team of AI agents to investigate questions together. Can they gain new capabilities through research and use them on unfamiliar problems, or does improvement reach a plateau?
Can research experience lead to new capabilities?
AI agents are software that can search sources, use tools and work through research tasks. We are designing a team that can divide those tasks, compare findings and check each other's reasoning. People help set questions, develop explanations and evaluate the results.
The Public Knowledge Layer is both shared context for the agents and the destination for checked findings. Gonka's network economics and AI hardware are the first topics. Each study is judged on its findings; whether it also develops transferable ways of working is a separate research question.
We distinguish efficiency on familiar tasks, transfer to a new field and capabilities that make previously unsuccessful investigations possible. Repeated trials distinguish accumulating gains from a plateau. Changes may be in the AI-agent team's organisation, tools and reusable skills rather than the underlying AI model.
The comparison protocol includes a strong single agent and a fixed-workflow team, with matched source access, task scope and resource budgets. Model upgrades and human contributions must be recorded separately from improvements learned by the agents. The research method explains the comparisons; they have not yet been completed.
What the AI-agent team may change
- roles and division of labour
- search and extraction tools
- memory and reusable skills
- how tasks are assigned and conflicting findings checked
- verification workflows
What the agents may not change on their own
- evidence requirements and acceptance criteria
- permissions to use sources or publish private material
- how people can correct the work or stop a run
- which decisions the agents are allowed to make
Have a source, contradiction or correction?