Show AI progress when it represents a real operation that matters to the user, such as retrieving approved documents or validating a proposed CRM update. Show sources when people need to inspect evidence. Do not present visible effort as proof that an answer is correct.
AI transparency can mean several different disclosures: a status indicator, an activity log, a source list, or a generated reasoning summary. They answer different user questions and should not be presented as interchangeable access to a model’s private internal process.
This guide connects behavioral research with practical workflow design. The examples are proposed operating patterns, not vendor-specific integrations, schemas, or guaranteed product features.
Show real work that helps a user evaluate progress, but never use a display of effort as a substitute for evidence of quality.
Should an AI system show its work?
Often, but first decide what the user needs to do. A status indicator helps someone understand whether a task is queued, running, waiting for review, or complete. An activity log records operations the system performed. Citations identify evidence attached to an answer. A reasoning summary offers a generated explanation, not a complete record of a model’s internal process.
Use a disclosure when it helps someone act. A long-running research task may need meaningful milestones. A factual answer may need source links and validation results. A proposed CRM change may need a review state and evidence reference. If a status cannot be tied to a verifiable event, omit it rather than adding narration that merely signals effort.
What the labor-illusion research does and does not show
Ryan W. Buell and Michael I. Norton studied operational transparency in their 2011 Management Science paper, “The Labor Illusion: How Operational Transparency Increases Perceived Value.” Operational transparency means showing aspects of the work performed while a service is being delivered.
In Experiment 1, 266 participants used a simulated online travel site. They experienced an instantaneous result or a wait of 10 to 60 seconds. During the wait, the site either showed a progress bar or displayed changing search activity, including sites searched and fares being compiled. Both versions returned identical itineraries and prices. Showing aspects of the operation increased perceived service value compared with a visually blank wait in the conditions tested.
Experiment 2 included 118 participants who compared instantaneous service with transparent or blind waits of 30 or 60 seconds. In the transparent condition, 62% preferred waiting 30 seconds and 63% preferred waiting 60 seconds over instantaneous service. These findings describe preferences in simulated service settings. They do not establish improved factual accuracy, conversion, or the same effect in every AI product.
Read the original paper on operational transparency.
A transparent wait can be preferred to an instant result in the tested setting. Apply that finding by exposing real milestones such as “sources retrieved” or “validation pending,” not by inventing activity or treating perceived effort as quality evidence.
Progress, sources, and reasoning summaries are different disclosures
Product interfaces vary by provider, model, release, interface, and configuration. The examples below describe specific documented capabilities, not a universal industry feature set.
| Disclosure | What it tells the user | Example |
|---|---|---|
| Progress | Current state or completion | “Validating response” |
| Activity log | Recorded system operations | “Repository search completed” |
| Sources | Evidence attached to an answer | Document title and link |
| Reasoning summary | Generated explanation or summary | “Summary of reasoning” |
Anthropic described visible extended thinking for Claude 3.7 Sonnet as a research preview. It listed helping users understand and check answers, supporting alignment research, and making the process interesting to observe as potential benefits. Anthropic also cautioned that displayed intermediate reasoning can be incorrect or incomplete and may not fully represent the factors influencing the model’s behavior. Read Anthropic’s announcement and limitations.
Google described Gemini Deep Research showing thoughts while browsing in a particular product experience. OpenAI’s API documentation describes optional reasoning summaries for supported models and says raw reasoning tokens are not exposed. OpenAI also described citations and a summary of the system’s thinking in its deep-research announcement. These disclosures are product-specific and should not be treated as identical behavior across providers. Read Google’s Gemini Deep Research announcement, OpenAI’s reasoning documentation, and OpenAI’s deep-research announcement.
Label information according to what it is: “Searching approved sources” for a real retrieval operation, “Sources used” for citations, and “Reasoning summary” for a generated summary. Do not label a generated explanation as a complete internal trace.
Design transparent activity around real events
Start with the operation, not the animation. Map the trigger, input, retrieval or tool actions, model task, validation, destination, and exception owner. Then expose only state changes that the system can verify. A long-running workflow might show “queued,” “retrieving,” “validating,” “awaiting review,” and “completed” or “failed.” Each displayed state should have a recorded event behind it.
For research tasks, preserve a source title, URL or document identifier, and retrieval time when available and permitted. Keep source records separate from the status log. A status event describes what happened in a run; a citation record identifies evidence attached to that run. Record validation outcomes separately so users can distinguish “retrieval completed” from “claims checked.”
Teams planning an operational agent can explore AI agent design and implementation. The appropriate activity display depends on the operations the system actually performs and records.
Three hypothetical patterns for transparent AI operations
Each pattern below uses the same control sequence: a real trigger produces bounded input, AI returns a constrained result, validation gates the next action, and an identified owner handles exceptions. The system of record remains responsible for authoritative data.
| Trigger and input | AI job and output | Validation gate | Action and fallback |
|---|---|---|---|
| Research question plus approved repository scope | Summarize retrieved documents, identify passages, and flag conflicts. Return a run ID, summary, and source references. | Confirm each source was retrieved, resolves to the intended document, matches the schema, and is authorized for the user. | Show the answer and source list. Route missing, inaccessible, or contradictory sources to research operations. |
| CRM record change plus approved text and record ID | Propose one allowed classification value with an evidence reference. | Check record identity, required fields, data type, allowed values, evidence, and policy. | Create a draft update or review item. Do not write back unresolved or contradictory values; the CRM process owner handles them. |
| Permitted document enters a processing queue | Extract or classify bounded fields and return structured output. | Check required fields, types, enumerated values, permissions, and deterministic business rules. | Complete only after required checks pass. Send incomplete extraction, timeout, or permission failure to a named reviewer or workflow owner. |
1. Research assistant with source retrieval
The sequence is: user submits a question, a retrieval service searches an approved repository, the AI summarizes retrieved material and flags possible conflicts, schema and source checks run, and the answer and source list appear in the interface.
Store a request ID, run ID, status events, output reference, and separate citation records, subject to access and retention rules. A citation should identify the source attached to that run, not simply repeat the prompt or date. Show a source only when it was retrieved for that run and the user is authorized to see it. If retrieval fails or a source does not resolve, show the failure and route the request to the research owner.
2. AI-proposed CRM classification
A record change can trigger a bounded classification request containing only the approved text and record identifier. The AI proposes one value from an allowed set and returns an evidence reference. Deterministic checks then validate the structure, record identity, required fields, allowed value, and applicable policy. The result becomes a draft field update or review item, while the CRM remains the system of record.
Use ordinary rules for explicit conditions such as a fixed status, threshold, required field, or approved market list. Reserve AI for ambiguous text or evidence comparison. For example, a proposal with a missing record match, unsupported value, or contradictory evidence should be rejected into an exception queue rather than resolved by guessing. The reviewer identity, decision, timestamp, and amended value should be stored separately from the original proposal.
Teams reviewing this kind of process can learn more about CRM systems and workflow design. This link describes a service area, not a particular CRM connector or approval feature.
3. Long-running document workflow
A permitted document enters a queue, extraction and classification run, deterministic checks validate required fields and allowed values, and the workflow either completes or routes an unresolved item to review. The document system or agreed records platform remains authoritative.
For incomplete extraction, show “awaiting review,” retain the output reference, and assign the item to a named reviewer. For timeout or permission errors, record a failed state and let the workflow owner investigate or retry. Do not show “completed” until required validation has passed.
The following event object is an illustrative schema, not a vendor-defined format. One event represents one status change in one execution.
{
"tenant_id": "tenant_illustrative_4",
"request_id": "req_illustrative_205",
"run_id": "run_illustrative_03",
"execution_attempt": 1,
"event_sequence": 3,
"event_type": "validation_completed",
"event_timestamp": "illustrative timestamp",
"status": "awaiting_review",
"validation_status": "passed",
"output_reference": "result_illustrative_31"
}
Set validation, ownership, and privacy rules before launch
Validate structured output before downstream use. Parse it against the expected schema, check required fields and types, enforce allowed values, and verify that cited sources resolve to the intended material. Apply business policy after structural validation. For high-impact changes, keep the AI result as a proposal until a person approves it or an explicitly approved policy permits automation.
Where retention policy permits, preserve the original input reference, model and prompt version, run identifier, output, validation result, and review decision. Limit displayed and retained document details to what each user is authorized to access. Apply the same access and retention rules to source material and activity logs.
Define data grain before storing records. One run identifies one execution. One event identifies one status change in that run. One citation identifies one source attached to a run. A prompt observation identifies one measured execution, while an aggregate summarizes a defined period and grouping.
Prevent duplicates with database-enforced uniqueness and transactional upsert where supported. For an event log, a candidate key is tenant ID plus run ID plus event sequence. For citations, use tenant ID plus run ID plus citation ordinal, or another key that matches the citation grain. A URL alone is not sufficient because one source may appear in multiple runs. A lookup followed by create can race when concurrent workers process the same request.
Keep aggregate metrics at their own scope. A share-of-voice or similar summary must retain dimensions such as period, engine, model, prompt set, geography, and metric definition. It should not be attached to one prompt execution as though it were an observed property of that run.
Measure clarity separately from accuracy
Evaluate the interface and the underlying task as separate questions. For the interface, measure task completion, time to a useful result, abandonment, clarification requests, and whether users understand what the system did. For system quality, measure citation validity, unsupported-claim rate, schema-validation failures, correction rate, and reviewer overturn rate.
When testing progress displays, keep the underlying operation and answer quality as comparable as possible. A longer wait or a higher perceived-effort rating is not evidence of a more accurate result. Also record the observation grain so individual events are not confused with prompt-level results or aggregates.
A practical rule for deciding what to reveal
Use transparency to make an operation understandable and its evidence inspectable. Show real, relevant activity; identify sources; label generated summaries accurately; validate outcomes independently; and make responsibility for exceptions clear.
- Does every displayed status correspond to a recorded event?
- Can an authorized user inspect the evidence the answer refers to?
- Are validation and approval stored separately from progress?
- Do source records, logs, and retained fields follow access and retention rules?
- Is a named owner responsible for failures, ambiguity, and unresolved cases?
