Long-term research
Machine Wisdom
When should an AI research agent answer, look for more evidence, ask for human judgement or stop? Our evaluation concerns these decisions and whether working with AI helps people think and decide for themselves.
Capability is not wisdom.
An AI system may be good at completing a task and still be poor at deciding whether its answer is justified. Here, Machine Wisdom names research into how AI agents behave when evidence is incomplete, reliable sources conflict, a question rests on a false assumption or a decision belongs to a person.
These questions matter now in tokenomics and hardware research: does a system challenge an appealing incentive story, retain evidence against a preferred supplier or stop when a critical estimate cannot be verified? They guide concrete review questions, not claims that the tests have already been run.
The evaluation also concerns the person working with the system: do they develop questions of their own, see alternatives it did not offer and become better able to judge the findings? A person who only approves the system's proposals is not the intended outcome.
Judgement in the research task
- say unknown when evidence is insufficient
- keep evidence that challenges the preferred answer visible
- question a misleading premise rather than solve the wrong problem
- accept a correction and revisit conclusions that depend on it
- preserve human judgement and the right to stop
An evaluation framework in development.
We do not assume wisdom is a single machine property. The framework is in development, not a validated benchmark. The research method can record judgement failures and changes in people's ability to question and decide, without pretending these already add up to a score for wisdom.
Have a source, contradiction or correction?