play icon for videos

AI for Social Good: Practical Uses, Risks and a First-Pilot Plan

Explore AI for social good across sectors. Choose a reviewable first use, assess reliability and governance, and measure operational value and real outcomes.

Updated
September 15, 2026
360 feedback training evaluation
Use Case
Impact & ESG portfolios · Practical guide

AI for Social Good: Practical Uses, Risks and a First-Pilot Plan

Explore AI for social good across sectors. Choose a reviewable first use, assess reliability and governance, and measure operational value and real outcomes.

Read the guide ↓

What is AI for social good?

AI for social good means using artificial intelligence to support beneficial social or environmental outcomes. The field includes education, accessibility, public services, environmental work, humanitarian response and community development. It involves businesses, research institutions, public agencies, associations and nonprofits.

The label describes an intention, not proof of benefit. A useful project needs a specific problem, a suitable task for AI and evidence that the resulting service or decision improves. It also needs a way to recognize and address harm.

ITU's AI for Good initiative takes this broad view of practical applications supporting sustainable development. This page helps an operational team choose a manageable starting point. The companion AI for social impact guide focuses on evidence, measurement and evaluation.

Where AI may help

The examples below are application areas, not claims that every product can deliver them or that a particular organization has achieved the outcome.

Scroll horizontally to see all columns →

AreaPossible taskWhat would need evaluation
Education and learningOrganize learner feedback or support suitable learning materialsAccuracy, accessibility, educational usefulness and learner safeguards
AccessibilityAssist with captions, translation or adaptation of informationMeaning, usability and quality for the intended audience
Environmental workHelp analyze relevant environmental observations or reportsData suitability, model performance and the action informed
Public and community servicesOrganize consultation responses and service feedbackCoverage, interpretation, privacy and follow-through
Health-service experienceReview appropriately collected feedback about service accessConfidentiality and usefulness; separate from clinical decisions
Employment and trainingReview barriers to applying learning or accessing opportunitiesAppropriate measures and human interpretation
Grant and portfolio workExtract report fields and organize evidence for reviewSource accuracy, missingness and human decisions
Humanitarian operationsHelp organize authorized field informationContext, reliability and the consequences of errors

AI is broader than a chatbot. Different tasks may involve classification, forecasting, computer vision, optimization or language models. The appropriate method depends on the problem and the consequences of getting it wrong.

A practical starting point for a growing organization

Start with a repeated task whose output can be checked before it affects people. For many teams, that is organizing existing evidence: survey comments, partner reports, member returns or program feedback.

This work is often constrained by staff time. A small team can collect more material than it can consistently review. The aim is to reduce the repeated effort of sorting and joining evidence while keeping the interpretation and decision accountable.

A good first task has a clear input, a defined output, known reviewers and a way to compare the new workflow with the current one. “Use AI to improve society” is too broad to evaluate. “Prepare a source-linked review of recurring barriers in this authorized feedback collection” is concrete enough to test.

Example: a professional network reviews member feedback

Fictional teaching example. A network receives 2,000 comments from member organizations across several regions. The team wants to understand obstacles to accessing a training opportunity. It already has relevant registration context and a defined collection period.

  1. Clarify the review question. Which reported barriers should the program team investigate?
  2. Define the themes. Reviewers agree what counts as scheduling, information, accessibility or opportunity to participate, with examples and an uncertainty category.
  3. Apply and review. AI assists with applying the definitions; reviewers check appropriate samples, difficult cases and contradictory accounts.
  4. Compare with context. The team examines approved group-level patterns alongside participation information, without treating the responding sample as all members.
  5. Choose and review an action. Staff select an improvement, return the findings and later examine whether the change helped.

Processing 2,000 comments is an output of the analysis workflow. A useful program change is a different result. Improved access is a further outcome that still needs evidence.

Measure the value at three levels

Scroll horizontally to see all columns →

LevelQuestionPossible evidence
Task performanceDoes the system do the assigned task adequately?Reviewed outputs, errors, omissions and performance across relevant cases
Operational valueDoes the full workflow help the team?Preparation, review and correction time; usable coverage; maintainability
Social or environmental outcomeDoes the resulting change benefit the intended people or setting?Appropriate outcome evidence, affected perspectives and examination of other influences

Do not report the first level as proof of the third. Faster processing does not necessarily mean better decisions, and a plausible recommendation does not establish that an intervention worked.

Use a suitable comparison with the current process and record both benefits and harms. Include the people affected when deciding what success means. The impact measurement guide helps distinguish activities, outcomes and contribution.

Sources help verification; they do not guarantee correctness

A cited answer can still misread its source, omit contrary evidence or attach a source that does not support the claim. A correct summary of biased or incomplete data can still lead to a poor conclusion.

Check source support, coverage, definitions and calculations. Review whether the output answers the actual question. Keep a way to correct errors and inspect changes after a model, codebook or source update.

NIST's Generative AI Profile includes confabulation among the risks to manage. Adding retrieval or citations does not remove the need for evaluation and oversight.

Make responsibility part of the workflow

The NIST AI Risk Management Framework organizes risk management around governing, mapping, measuring and managing. For a small operational pilot, translate that into a few concrete decisions:

  • Who owns the purpose, definitions and acceptance criteria?
  • Which information may the system access, and for what use?
  • What errors or omissions would matter most?
  • Who reviews outputs and can stop or change the process?
  • How are corrections, incidents and version changes recorded?
  • How can affected people raise concerns or challenge an outcome?

Human oversight needs time, authority and a workable review interface. A person who only approves a fluent answer without checking it is not a strong control.

Match the safeguards to the consequences

Organizing a draft reading list and recommending an individual service decision are different uses. Funding, employment, eligibility, clinical care and safeguarding decisions can have serious consequences and require the appropriate professional, procedural and legal context.

Do not start by automating those decisions because they are time-consuming. Begin with clearly bounded assistance where errors can be identified and corrected before consequential action. Keep specialist systems and expertise where needed.

Protect data across collection, model access, summaries, exports and retention. Do not upload sensitive source material into a general tool without verifying the arrangements appropriate to that information and purpose.

Review whose experience is represented

AI cannot infer the experience of people who were never reached by collection. Examine coverage, language, accessibility and missing responses before interpreting the patterns.

Assess output quality across relevant groups and types of input, including unusual or difficult cases. A strong average can hide poor performance for a particular language or context. Avoid treating a single fairness statistic as proof that all harmful bias has been addressed.

Where the task involves open-ended accounts, keep uncertainty and disagreement visible. A theme is an interpretation, not an objective fact simply because it was assigned automatically.

Compare total effort, not only the first demonstration

Include the work to prepare data, configure access, define the task, test outputs, review exceptions, make corrections and maintain the workflow. A tool that produces a quick draft may still require substantial recurring checking.

For qualitative analysis, the recurring burden often includes reapplying a revised codebook and joining themes back to quantitative context. Sopact's approach keeps people in control of the definitions while reducing that repeated application and assembly work.

The qualitative and quantitative analysis guide shows the visual workflow and an explicitly illustrative effort comparison. Use your own volumes and review requirements to assess the likely benefit. There is no universal time-saving guarantee.

Plan a first pilot you can evaluate

  1. Choose one recurring problem. Keep the purpose and users specific.
  2. Document the current process. Include effort, quality and known gaps.
  3. Select appropriate data. Start with synthetic or suitably authorized examples.
  4. Define the review method. Include difficult cases, errors and omissions.
  5. Run the whole cycle. Test updates, corrections, access and handoff to a real decision.
  6. Decide what the evidence supports. Expand, revise or stop the use based on observed performance and consequences.

Small does not mean superficial. A narrow pilot can include a realistic definition change, a restricted source and an output that the system should decline to answer.

Where Sopact fits

Sopact focuses on connected operational evidence: recurring collection, relevant person or organization context, qualitative and quantitative analysis, and human review. This can serve associations, businesses, program providers, funders and service teams whose evidence outgrows disconnected forms and files.

The aim is a workflow the team can manage with clear definitions and controls, without repeatedly rebuilding data joins and reports. Verify the actual capabilities and safeguards your task requires rather than treating the general approach as proof of every product feature.

Explore AI document analysis, AI grant management and AI data collection for specific workflows. The Impact Measurement & Reporting course connects the evidence plan to useful reporting.

For communicating results, use the impact-report writing guide and report examples. State the actual benefit and uncertainty, not only the AI task completed.

Watch: connected collection and analysis

This introduction shows Sopact's operational evidence approach. It is one application within the broader AI-for-social-good field, not a demonstration of every domain above.

Frequently asked questions

Is AI for social good limited to nonprofits?

No. Businesses, associations, researchers, public agencies and nonprofits can use AI toward beneficial social or environmental outcomes. The intended benefit needs evidence regardless of the organization's type.

What is a manageable first use?

A bounded, reviewable task such as organizing authorized feedback or extracting defined fields from reports can be a practical starting point. Evaluate the complete workflow and consequences before expanding.

Do citations make an AI answer reliable?

They help reviewers check it, but do not guarantee that the source supports the claim or that the data is complete. Verify interpretation, coverage and calculations as well.

Can AI eliminate bias?

No system or single check guarantees that. Review data coverage, output performance across relevant cases, affected perspectives and the controls for correcting or challenging errors.

How do we know whether the project creates social good?

Measure task performance, operational value and the intended social or environmental outcome separately. Include harms and other influences; faster processing alone is not proof of social benefit.

Explore Impact Measurement →