Every year, billions of dollars in healthcare reimbursements depend on decisions made by medical auditors. Yet those decisions often require reviewing hundreds of pages of clinical documentation under immense time pressure!!
Designing Trust Between AI and Human Expertise
Reimagining Medical Claim Auditing with APEX AI Recommendation
Every medical claim tells a story. Not through a single document.But through hundreds of pages of physician notes, nursing observations, laboratory reports, imaging results, discharge summaries, medication history, and clinical documentation collected throughout a patient's hospital stay.
For a medical auditor, understanding that story isn't optional!!
It determines whether the correct Diagnosis Related Group (DRG), diagnosis codes, and procedure codes are assigned decisions that directly influence reimbursement, regulatory compliance, and ultimately millions of dollars in healthcare spending.
The challenge wasn't that auditors lacked expertise. The challenge was that they were expected to uncover critical evidence hidden across hundreds of pages while balancing accuracy, productivity, and compliance. That was the problem APEX set out to solve.
Home Screen - APEX AI Recommendation
Understanding the Problem
Before designing anything, I spent time understanding how medical auditors actually worked. Unlike traditional consumer applications, these users weren't browsing products or booking appointments. They were making high-stakes clinical decisions where even a single incorrect diagnosis code could affect reimbursement or trigger compliance issues.
Through stakeholder workshops, workflow walkthroughs, observation sessions, and discussions with medical coding experts, clinical auditors, product owners, and AI engineers, several recurring pain points emerged.
• A single claim could contain hundreds of pages of medical records spread across multiple document types.
• Auditors spent most of their day searching rather than analyzing.
Reviewing a recommendation meant jumping repeatedly between
• Medical Records
• Coding Systems
• Policy Documentation
• DRG References
• Audit Worksheets
Every interruption increased cognitive load.
AI Wasn't the Problem. Trust Was.
• Early AI models could already recommend coding changes. However, auditors repeatedly asked one question.
• However, auditors repeatedly asked one question.
Why is the AI recommending this?
• Without supporting evidence, recommendations felt like black boxes.
• No matter how accurate the AI became, auditors would never approve recommendations they couldn't validate.
• That insight fundamentally changed the direction of our design.
• We weren't designing an AI recommendation engine.We were designing an AI explanation experience.
Defining the Design Challenge
Our challenge became much broader than displaying AI output. We needed to design an experience that would allow auditors to
- Understand recommendations quickly
- Validate clinical evidence
- Maintain confidence in their own decisions
- Provide structured feedback to continuously improve the AI
- Complete the entire audit without feeling overwhelmed
AI should Accelerate Human Expertise. NOT replace it.
Research & Discovery
To ensure our design decisions reflected real-world workflows, I collaborated closely with cross-functional teams throughout discovery.
Performed Stakeholder Interviews with:
- Medical Coding Auditors
- Clinical Documentation Improvement (CDI) Specialists
- Product Managers
- AI & NLP Engineers
- Business Analysts
- Compliance Teams
Rather than asking what features they wanted, I focused on understanding how they made decisions.
Some of the questions we explored included:
- How do you validate an AI recommendation?
- What makes you reject a recommendation?
- Where do you spend the most time during an audit?
- What information do you need before changing a DRG?
- Which parts of the workflow create the most frustration?
Interestingly, almost nobody asked for "more AI." Instead, they wanted better visibility into why the AI reached its conclusion. That insight shaped almost every design decision that followed.
Key Insights
Our research revealed five important findings.
- Evidence builds trust
- Auditors trusted recommendations only after reviewing supporting documentation.
- Context should never be hidden
- Closing the medical record to inspect recommendations created unnecessary mental effort.
- Search consumed too much time
- Finding one diagnosis inside hundreds of pages was often slower than validating it.
- Every disagreement is valuable
- Rejected recommendations weren't failures.
- They represented opportunities to improve future AI models.
- Different users required different depths of investigation
- Some users only needed a summarized explanation.
- Others wanted to inspect the raw medical record line by line.
- The interface had to support both.
Translating Research into Design?
Design Decisions
Instead of designing isolated features, I designed an ecosystem where every component supported a specific user decision.
Decision 1
• Keep Recommendations and Medical Records Together
• Instead of opening separate screens, I designed a split-screen workspace where auditors could simultaneously review AI suggestions while reading medical records.
• This significantly reduced navigation and preserved context throughout the audit
Decision 2
Rather than simply displaying a new DRG or diagnosis code, every recommendation included
• Suggested Changes
• Supporting Rationale
• Policy Evidence
• Confidence Score
Decision 3
One design challenge involved communicating AI certainty.
• Displaying recommendations without context risked creating blind trust.
• Instead, we introduced confidence indicators that communicated how certain the AI was while still encouraging auditors to verify clinical evidence.
• The confidence score became a conversation starter rather than a final answer.
Decision 4
Most AI products collect feedback after the experience.
We wanted feedback to become part of the decision itself.
Auditors could provide structured feedback at three levels:
• Overall Claim
• Individual Recommendation
• Policy Criteria
This transformed every audit into a learning opportunity for future AI improvements.
Decision 5
Searching hundreds of pages manually wasn't scalable.
We introduced an AI Assistant capable of answering contextual questions such as:
• Why was the DRG changed?
• Where is BMI documented?
• What evidence supports this diagnosis?
Instead of forcing users to hunt for information, the assistant brought relevant evidence directly into the workflow.
Decision 6
While many auditors were satisfied with summarized explanations, experienced users often wanted to inspect every occurrence within the medical record.
Rather than overcrowding the primary interface, we integrated an NLP Viewer that enabled deeper exploration through:
• Keyword Search
• Entity Highlighting
• Section Navigation
• Page References
This preserved simplicity while still supporting expert-level investigation.
Working Within Constraints
Like most enterprise healthcare products, design decisions weren't driven by aesthetics alone.
Several constraints shaped the solution.
• Every recommendation required traceable clinical evidence.
• Transparency wasn't optional—it was a compliance requirement.
• Claims frequently exceeded hundreds of pages.
• Performance, navigation, and readability became critical UX considerations.
• Machine learning models continuously evolved.
• The interface had to gracefully accommodate recommendations with varying confidence levels without misleading users.
• Some auditors had decades of coding experience.
• Others were newer to the role.
• The interface needed to feel approachable without limiting advanced workflows.
Measuring Success
Rather than focusing solely on AI accuracy, success was measured through user outcomes.
The redesigned experience enabled auditors to:
- Review AI recommendations without losing clinical context.
- Validate coding decisions through evidence-backed explanations.
- Reduce time spent searching across lengthy medical records.
- Capture meaningful structured feedback to continuously improve AI recommendations.
- Maintain human control over every final coding decision.
What This Project Taught Me
This project fundamentally changed how I think about designing AI products. Before APEX, I believed the success of AI depended on how intelligent the model was. Working closely with medical auditors taught me something different.
The intelligence of an AI system matters far less than how confidently people can validate its decisions. Designing AI isn't about replacing experts. It's about reducing uncertainty. It's about helping professionals arrive at better decisions with greater confidence.
The most successful interface we created wasn't the chatbot, the recommendation engine, or even the search experience. It was the invisible layer of trust connecting human expertise with machine intelligence. That, more than anything else, became the real product.