AI for Altruism

AI for AltruismAI for AltruismAI for Altruism
  • Home
  • Insights
  • Contact
  • Team
  • Research
  • For Nonprofits
  • More
    • Home
    • Insights
    • Contact
    • Team
    • Research
    • For Nonprofits

AI for Altruism

AI for AltruismAI for AltruismAI for Altruism
Donate
  • Home
  • Insights
  • Contact
  • Team
  • Research
  • For Nonprofits
Donate

How to Evaluate an AI Tool

You probably already have one. AI features arrived inside the software your organization was already paying for: your email, your documents, your CRM, your note-taker. The useful question is no longer which tool to adopt. It is whether the one you are using is doing what you think it is doing.


This page is about how to find out. It is the same skill whether you are judging a vendor's sales claim, a funder's dashboard, or a tool your own team built.

An accuracy number is at least three claims

When someone tells you a system is 95% accurate, they have compressed three separate claims into one number, and any of the three can be the weak one.


What was actually measured. Accuracy at what task, on which examples? A tool that is 95% accurate at classifying clearly-worded English requests may be far worse on the requests your community actually sends you.


Who judged it. Someone or something decided what counted as a right answer. A vendor's own scoring, an automated grader, and two of your staff reading the same outputs can reach genuinely different verdicts on identical results.


Whether it holds in your setting. Performance measured on a vendor's data is a claim about the vendor's data. Your caseload, your languages, your intake forms, and your edge cases are not theirs.


Ask which of the three a given number covers. Most marketing numbers cover only the first, and only partly.

What this looked like when it happened to us

We ran a preregistered comparison of two ways of answering questions over a research library, and we scored the results two defensible ways.


A holistic rubric, the kind that produces the number on a slide, judged one system better grounded in its sources. A strict check that asked whether each individual citation actually supported the claim attached to it favored the other system, and not narrowly: its cited claims held up 40.2% of the time against 18.9%.


Same systems, same outputs, two reasonable scoring methods, opposite verdicts. Neither number was wrong. They measured different things, and only one of them was going to end up on a slide.


Two honest limits on that finding, because they are the kind we would want disclosed to us. The scoring was done by language models rather than human reviewers. And the strict check asked whether each claim followed from the evidence the system itself cited, which is not the same as asking whether it followed from the original underlying document.


The full study, the registration filed before we ran it, and the code are linked from our research page.

Six questions to put to a vendor

Ask these in writing. The quality of the answers tells you as much as the answers.

  1. What exactly did you measure, on how many examples, and where did those examples come from?
  2. Who or what decided which answers were correct, and did anyone independent check that judgment?
  3. What is the performance on the cases where it does worst, not on average?
  4. How does the tool behave when it does not know? Does it decline, or does it guess fluently?
  5. What happens to the data we put in? Is it used to train anything, who can see it, and can we get it back or have it deleted?
  6. What would we see if it started getting worse, and how quickly would we see it?

A vendor who cannot answer 1 and 2 has not measured their product; they have described it. A vendor who cannot answer 5 is not a vendor you can responsibly put beneficiary data through.

Checking it yourself, without a data team

You do not need a statistician to get a defensible answer. You need about half a day.


Write down what a good answer looks like, before you look at any output. This is the step that gets skipped, and skipping it is what makes an evaluation unfalsifiable. If the criteria are written afterward, they will quietly bend toward what the tool already does.


Pull a sample of real cases, not demo cases. Thirty to fifty is enough to see a pattern. Include the awkward ones on purpose: the long ones, the ones in another language, the ones with missing fields, the ones a new staff member would get wrong.


Have two people score them independently, then compare. Where your two scorers disagree, you have found something more valuable than a score. Either your criteria are ambiguous or the task is genuinely harder than it looked. In our own work, our human-to-automated agreement came in at 0.23 against the 0.60 we had committed to in advance, and that gap told us more than the headline result did.


Decide in advance what result would make you stop. A threshold chosen after seeing the numbers is not a threshold.

What to do with the answer

If it works, write down what you measured and when, and put a date on rechecking it. Tools change under you without notice, and a result from last year is a historical fact rather than a current one.


If it does not work, that is a finding and not a failure. It cost you half a day instead of a program year. We publish our own negative results for the same reason: the discipline only means anything if it survives the times it goes against you.


If you cannot tell, that is also a result, and usually it means the criteria were not specific enough to be checkable. Go back to the first step.

Where we can help

We are a research nonprofit rather than a vendor, so we have nothing to sell you here and our answer to "should we use AI for this?" is sometimes no. If you want a second pair of eyes on a vendor claim, or help designing an evaluation you can defend to your board, get in touch.

AI for Altruism (A4A) is a 501(c)(3) nonprofit organization.

Centennial, Colorado, USA

Copyright © 2026 AI for Altruism, Inc. - All Rights Reserved.

Please donate today.

  • Privacy Policy
  • Donation Policy
  • For Sponsors
  • Insights

This website uses cookies.

We use cookies to analyze website traffic and optimize your website experience. By accepting our use of cookies, your data will be aggregated with all other user data.

Accept