All Insights

Using Synthetic Data to Build Better AI Workflows in Fundraising

By Ryan Clement | Ryan Clement Consulting · Published · Updated

One of the biggest questions nonprofit professionals have about AI is also one of the right questions to ask:

What happens to our donor data?

Fundraising data can include names, addresses, giving history, wealth indicators, relationship notes, family information, employment details, engagement history, and other information that should not be treated casually.

That does not mean nonprofits have to avoid AI.

It means we need better ways to test what AI can do before exposing real constituent information unnecessarily.

One of the approaches I use regularly is synthetic data.

What synthetic data actually means

Synthetic data is artificially created data designed to reproduce the kinds of patterns you want to study without representing actual people.

For prospect development, that might mean building a fictional portfolio containing records such as:

  • A high-capacity prospect with little engagement
  • A loyal donor whose giving has declined
  • Someone who has remained in qualification too long
  • A prospect with an open opportunity but no recent contact
  • A highly engaged constituent without a defined next step
  • Someone with strong wealth indicators but limited philanthropic evidence

None of those people needs to exist.

The records can be fictional while the fundraising problem remains completely real.

That distinction makes synthetic data particularly useful for AI development.

NIST describes synthetic data as newly generated records intended to preserve useful properties of an underlying population while representing artificial individuals rather than the original people.

Why I use synthetic data when testing AI models

When I am evaluating a new scoring model, prioritization method, prompt, or workflow, I want to know how the system behaves.

Does it overvalue wealth?

Does it recognize sustained engagement?

What happens when data is missing?

Does it confuse capacity with readiness?

Can it explain why one prospect received greater priority than another?

Can I intentionally create an edge case and see whether the system handles it correctly?

Those are model-development questions.

I do not need someone's actual donor record to answer them.

Synthetic data gives me room to build, test, break, revise, and test again.

That is especially useful for something like portfolio prioritization. A synthetic dataset can contain hundreds of fictional prospects with different combinations of capacity, affinity, giving history, engagement, stage, opportunity activity, and recent contact.

Then I can see whether the model behaves the way an experienced prospect-development professional would expect.

Synthetic data is also useful for demonstrations and training

There is another advantage: it makes AI easier to teach.

When demonstrating a new workflow to a fundraiser, advancement leader, researcher, or board member, there is usually no reason that the demonstration needs to contain an actual donor.

A synthetic portfolio can look and behave like real advancement data without putting someone's personal information on a screen.

That means teams can learn how the system works before discussing whether—and under what controls—it should ever interact with production data.

A useful development sequence is:

Define the problem → Build synthetic data → Test the workflow → Find failure points → Adjust the system → Establish governance → Consider approved real-world use

De-identification is another useful tool, but it is different

Sometimes you do need to analyze patterns from an existing dataset.

That does not automatically mean every identifying field has to travel with it.

For many analytical workflows, the AI does not need:

John Smith
123 Main Street
Cleveland, Ohio

It may only need:

Constituent ID 004872

along with the fields actually required for the analysis.

The organization's own systems can maintain the connection between the internal identifier and the person.

The AI workflow does not necessarily need to know who that person is.

This is essentially data minimization: provide the system with the information required for the task and avoid providing information it does not need.

NIST identifies removal of direct identifiers and transformation of other identifying fields as established de-identification techniques, while also warning that simple masking is not automatically sufficient protection.

That caveat matters.

Replacing a name with an ID number does not automatically make a dataset anonymous.

A combination of employer, ZIP code, exact gift amount, board membership, age, property information, and other fields could still make a person identifiable.

That is why the real question should always be:

What information does this workflow actually need?

Synthetic data and de-identified data are not the same thing

This distinction is important for nonprofit teams adopting AI.

Synthetic data creates fictional records designed to represent useful patterns.

De-identified or pseudonymized data begins with real records and removes or transforms identifying information.

They solve different problems.

For early experimentation, model testing, demonstrations, training, and workflow development, I generally prefer synthetic data whenever possible.

There is simply less reason to expose real donor information while you are still figuring out whether the workflow works.

When real organizational data eventually becomes necessary, that should happen within the organization's approved technology, privacy, security, and data-governance practices.

Synthetic does not automatically mean risk-free

There is an important limitation.

Not every synthetic dataset provides the same level of privacy protection.

If synthetic data is generated directly from sensitive source data, poorly designed methods can preserve enough information that some characteristics of the original individuals could potentially be inferred.

NIST has specifically warned that some synthetic-data techniques without formal privacy protections can remain vulnerable to privacy attacks.

So I would not treat the word synthetic as a compliance checkbox.

Organizations should still understand:

  • How the synthetic dataset was generated
  • Whether real constituent information was used to create it
  • What information the resulting records preserve
  • Whether individual records could potentially be reconstructed or inferred
  • Where the data is stored
  • Who has access
  • Whether the AI environment has been approved by the organization

Good governance still matters.

This is ultimately about designing the workflow correctly

A well-designed AI workflow should not begin with:

“Give the model everything we have.”

It should begin with:

“What is the minimum information necessary to solve this problem?”

That might mean synthetic data.

It might mean internal IDs instead of names.

It might mean removing addresses and personal details.

It might mean aggregating certain fields.

And sometimes it means deciding that particular information should not enter the AI workflow at all.

Responsible AI use in nonprofit work should include asking:

  • Is this information confidential?
  • Has the organization approved this AI tool?
  • Am I allowed to enter this type of information?
  • Can names and identifying details be removed?
  • Can the situation be described without exposing private data?
  • Is AI supporting professional judgment rather than replacing it?

Those questions should come before the prompt.

The better question for nonprofit AI

The conversation around AI and donor data often becomes:

“Can we put our fundraising data into AI?”

I think there is a better first question:

Can we prove the workflow works before we ever give it real donor data?

In many cases, the answer is yes.

Build a synthetic portfolio.

Create the difficult scenarios intentionally.

Test the model.

See where it fails.

Change the logic.

Run it again.

Make sure a professional can understand and challenge the results.

Then decide what information the production version actually requires.

Responsible AI does not start with restricting what people can build. It starts with designing better ways to build it.

That is how nonprofits can experiment with AI while continuing to respect the people represented by their data.