Case Study · Transcend · August 2026
When Trust Is the Product: Building AI to Strengthen Expert Work
How AI can increase productivity while keeping human expertise and judgment at the center

Challenge
UN offices at the frontlines of peace, security, mediation, and political affairs were synthesizing decades of documentation across six to seven separate sources per question — slow, fragmented work in a setting where speed and accuracy carry consequence.
Solution
Transcend designed an AI agent grounded in official UN documentation, harnessed in diplomatic practice and strict analytical standards, then validated it through a moderated study with senior practitioners.
Results
Up to 50% research-time savings, 80% task success, 80% adoption intent, a 3.8/5 user-reported trust rating, and advancement to Phase 2 — scaled to support up to 300 users.
When a UN political affairs officer sits down to write a one-page briefing on a peace process, the obstacle is rarely a lack of material. It is the opposite. Decades of resolutions, envoy reports, press statements, and field consultations sit scattered across disconnected sources, each demanding its own search, its own verification, its own judgment call about what to trust. Coordination and validation consume hours in a setting where every hour, and every error, carries consequence.
Transcend was engaged to design and validate an institutional AI agent grounded in official UN documentation —retrieval-augmented generation with in-text citations and source traceability— and harnessed in diplomatic practice and strict analytical standards. In other words, purpose-built to deliver the acceleration of modern AI without the risks that make general-purpose AI tools unsuitable for sensitive work.
In institutions where mistakes are measured in consequence rather than convenience, trust and transparency are requirements in an AI system.
The Approach: Validation Before Scale
Transcend piloted the AI agent with 15 users and conducted a moderated usability and pilot validation study with a subset of senior UN professionals spanning political affairs, mediation, and knowledge management. The research team used a think-aloud protocol across scenarios drawn from real practice: drafting a one-page briefing, tracing a citation back to its source, filtering by document type, etc. The aim was not simply to confirm that the platform functioned. It was to determine whether the platform earned the confidence of the practitioners whose judgment it exists to support.
The Results
The platform cleared four of six feasibility benchmarks, and where it fell short, it pointed with unusual clarity to exactly what remained to be built.
Grounding the system in official documentation, and pairing that grounding with visible in-text citations, changed how participants related to the output in front of them — from the reflexive skepticism most brought in from general purpose Large Language Models (LLM), to something closer to conditional confidence.
Practitioners trusted this system in a way they did not trust general-purpose tools for institutional work.
Two scores, trust rating and citation confidence, came in slightly under the study's 4-out-of-5 target. Neither reflected a weakness in the agent as the harness and infrastructure performed precisely as designed and to expectations. Rather, both traced to a single, addressable cause: the pilot corpus surfaced too few sources per answer, and too narrow a range of them. Participants said, consistently and without prompting, that they would extend more trust the moment that changed.
The Insight: Trust Requires Evidence and Accuracy
The study's most important finding was also its most transferable: in institutional environments, trust is earned through transparency — how many sources support a claim, how diverse those sources are, and how easily a user can trace an answer back to its logic to confirm accuracy. Where citation depth was thin, even accurate answers drew scrutiny. Where the system was honest about the limits of its own coverage, users responded with greater, not lesser, confidence. Credibility, in other words, is a design problem, and evidence is the interface.
A second, more behavioral finding has already shaped the Transcend platform. Users arrive at any new AI tool carrying the mental model that general-purpose LLMs or chatbots have made them accustomed to: type a prompt, get an answer, refine by typing more. The behavior is applied even to interfaces explicitly designed to work differently. Controls that fight that instinct simply go unused. 60% of participants never noticed data filtering controls, defaulting instead to refining their queries through the prompt box itself. This sent a corrective signal to the Transcend design and engineering team to solve for data transparency and selection differently.
Participants also pointed to concrete, addressable gaps: deeper citation coverage, clearer visibility into which documents the system could draw from, and more transparency about why some questions were unanswered or redirected. Each of these fed directly into optimizing Transcend's platform.
The Outcome and What's Next
The pilot met its feasibility threshold and advanced into a second phase with a focused roadmap: surface source coverage explicitly, deepen citations, align the interaction model with how people actually work, expand the recency of the knowledge corpus, and scale to support up to 300 users. These deliberate, evidence-led refinements turn a capable system into a trusted institutional one.
The UN's willingness to pilot a capability like this makes it a first mover among multilateral institutions, and a proving ground for how AI can responsibly support diplomatic and peacebuilding work. The combination of early adoption and disciplined scrutiny is itself a model other institutions can draw on.
Transcend provides an analysis and decision support platform for institutions that cannot afford to be wrong: governments, multilateral bodies, and organizations operating in complex geopolitical and conflict environments. This engagement is a demonstration of our method — design grounded in domain expertise, validated with the experts doing the real work, and engineered so that trust, traceability, and human judgment remain at the center of the build, not bolted on after the fact.
If your organization is exploring how to deploy AI responsibly in a high-stakes environment, we welcome a conversation.