Research method

Research Method

Our method for investigating questions, checking findings and building on earlier research with people and AI agents. The first comparison protocol pairs agent-produced work with research drafts already prepared by our team.

Status
In development
Page updated
2026-09-05
Next output
Detailed test plan for the first research comparison

Research that builds on earlier work.

We are developing a shared research library and a team of AI agents: software that can search sources, use tools and investigate questions. In this research model, people and agents draw on shared knowledge and return checked findings for others to build on.

Gonka's network economics and AI hardware are the first investigations in an open-ended agenda. The value of this research does not depend on an AI-agent team outperforming a simpler approach.

Our team has existing human-directed drafts on these topics, which give the first trials something concrete to compare against. The research agenda extends to questions without a prior report. The library, agent team and test plan are in development; results from trials combining them are still ahead.

People help shape the investigation.

People and AI agents can both propose questions and challenge the assumptions behind them. People retain responsibility for what to pursue and what to do with the findings. Human work includes forming hypotheses and developing explanations, not only checking an answer at the end.

Selected results require external domain review, with disagreements between reviewers and the research team recorded. Where quality depends on interpretation, the review must state the reasoning behind its judgement.

What did the research miss, add or fail to support?

Give each research approach the same question and scope, and specify which sources are available up to what date. Record the AI models and tools used, computing costs and human effort, including question-setting, analysis, review and interventions. Assess missing findings, new discoveries and unsupported claims against evidence; agreement with an earlier report is not proof of correctness. Where effort on a prior draft was not recorded, report that gap rather than inventing a cost comparison.

Compare agents working with and without the research library to test the value of reuse. Compare an AI-agent team that can change how it works with a team following a fixed workflow to test adaptation. Include a strong single agent with the same source access and a comparable budget: using several agents must earn its cost. Report cost and quality separately, and distinguish gains from a model upgrade or human intervention from improvements learned by the agents.

What does each check establish?

A reproducible simulation may still model the wrong market. A source can support a fact without supporting the conclusion drawn from it. We keep three judgements distinct:

  • Reproducibility: can someone repeat the calculation or experiment and obtain the reported result?
  • Empirical support: do observations support the assumptions and claims about the world?
  • Interpretation: how well does the explanation account for the evidence and competing explanations? Expert disagreement belongs in the assessment.

Keep earlier reports out of independent research runs.

The earlier report is a reference for comparison, not an answer key. Record how it was produced, including AI assistance. Its conclusions remain open to challenge.

Agents attempting the same investigation must not see that report or summaries derived from it, whether through the research library, agent memory or search results. They may use the underlying sources allowed by the test plan. Check for accidental exposure to the prior answer, disclose risks that cannot be ruled out and compare the reports only after the run is complete.

Reusing prior reports to answer new questions is a separate experiment. It tests the value of accumulated context, not independent reproduction, and must be labelled accordingly.

Review before returning findings to shared knowledge.

Our method is to gather sources, extract and challenge claims, develop explanations, evaluate results and review them before publication. Only reviewed findings enter the shared library, with their evidence and uncertainty preserved. A correction must also update conclusions that relied on the error; repeating an unchecked claim does not make it accepted knowledge.

The AI-agent team may change its roles, tools, memory and workflow in response to observed failures. It may not change the standards used to judge its results or give itself new decision-making powers. Machine Wisdom asks whether the agents retain conflicting evidence, question misleading assumptions and help readers exercise their own judgement; its assessment framework remains in development.

Have a source, contradiction or correction?