AI for Altruism

AI for AltruismAI for AltruismAI for Altruism
  • Home
  • Insights
  • Contact
  • Team
  • Research
  • For Nonprofits
  • More
    • Home
    • Insights
    • Contact
    • Team
    • Research
    • For Nonprofits

AI for Altruism

AI for AltruismAI for AltruismAI for Altruism
Donate
  • Home
  • Insights
  • Contact
  • Team
  • Research
  • For Nonprofits
Donate

Research

We register our studies before we run them, and we publish the result either way. Every study below links to the registration filed before the data existed, the preprint, and the code and data needed to check the result independently.

What we predicted, and what actually happened

We lead with the three occasions our own hypothesis did not survive the measurement, because how an organization reports its failures is the most useful thing you can know about how it reports everything else.
 

We predicted a compiled wiki would cost less per query than retrieval. It cost 21× more.

 

Our registered hypothesis was that compiling a research corpus into a maintained wiki would pay for itself across many queries. Measured, the wiki spent about 21× more per query than vector retrieval. The amortization we had predicted is not merely unsupported but unavailable: the break-even calculation returns a negative number of queries. We made that reversal the paper's headline rather than a limitation.

· Preprint: arXiv:2605.18490 

· Registration: osf.io/zemhp 

· Code: github.com/ai4altruism/rag2compare
 

We set a reliability threshold in advance, and our own grading did not clear it.
 

Our progressive-disclosure ablation registered a Cohen's κ of 0.60 as the required agreement between human and automated grading. The pooled result was 0.23. It appears in the paper's abstract rather than its appendix, together with the condition under which our non-inferiority claim does not hold.

· Preprint: arXiv:2607.04576 

· Registration: osf.io/feka7 

· Materials: osf.io/a53mz
 

We tested whether context prompts make transcription more accurate. They did not.
 

A preregistered ablation on archival oral-history audio found no detectable change in aggregate word error rate from prompt-level context, a null result against our own working assumption. The effect that does exist is narrower and invisible to the headline metric: recall of specific listed terms improves while aggregate error stays flat. That is a finding about the measure as much as about the method, and it changed what we build.
· Registration: osf.io/ns49b 

· Materials: osf.io/9nsvr 

· Preprint in preparation
 

Why we work this way

AI systems are increasingly evaluated by scores whose validity is rarely argued explicitly. Our research is about that gap: whether a measurement supports the claim being made on it, and what it takes to know. An organization making that argument has to be able to survive the same test, so we register the analysis before we look at the data, publish results that go against us, and ship the artifacts that let someone else disagree with us on the evidence.

Check it yourself

Our code and data are public. The reproduction kit for the retrieval comparison is at github.com/ai4altruism/rag2compare; other tools and infrastructure are at github.com/ai4altruism. 

Registrations and materials are deposited on OSF and linked with each study above.

Also on the record

Patent application. System and Method for Progressive Learning Using Iterative Assessment and Generation. U.S. Patent Application #18/964,674, filed December 2024.

AI for Altruism (A4A) is a 501(c)(3) nonprofit organization.

Centennial, Colorado, USA

Copyright © 2026 AI for Altruism, Inc. - All Rights Reserved.

Please donate today.

  • Privacy Policy
  • Donation Policy
  • For Sponsors
  • Insights

This website uses cookies.

We use cookies to analyze website traffic and optimize your website experience. By accepting our use of cookies, your data will be aggregated with all other user data.

Accept