
Key results
- Massive scale processing
Successfully automated extraction across 200,000+ referral documents, eliminating thousands of hours of manual data entry work.
- Format diversity mastery
Developed a system capable of handling 500+ different document formats from various healthcare providers, each with unique layouts and structures.
- Strong generalization
Demonstrated 75% accuracy on 425 entirely new, unseen document types, proving the system's ability to adapt to evolving formats and sources.
The document processing bottleneck
The company processed a large volume of referral documents submitted by hospitals, clinics, and healthcare providers. These documents varied widely in structure, quality, and format. Some were digitally generated, others scanned or faxed, and many lacked consistent labeling. Critical patient identifiers such as names, dates of birth, addresses, and insurance details appeared in different locations across documents.
Intake teams manually reviewed each document, searched for required information, and entered the data into internal systems to locate or create patient records. This process was slow and prone to errors, particularly when dealing with unfamiliar formats or unclear scans. As referral volumes grew into the hundreds of thousands, the manual approach struggled to scale. A single misread number or transposed digit could delay equipment delivery or create duplicate records. The problem wasn’t just speed but maintaining accuracy under increasing pressure.
Designing an automated extraction workflow
To address this challenge, the company worked with Data Science Dojo to build an automated system capable of extracting HIPAA identifiers directly from referral documents. The objective was to reduce manual effort while improving consistency in patient lookup and record creation.
The project began with a detailed review of the document landscape. Over two hundred thousand referral documents were analyzed, revealing more than five hundred distinct formats with no standardized structure. This analysis helped establish a representative document corpus that reflected real-world variation rather than idealized templates. The diversity was significant: some documents had identifiers at the top, others buried in paragraphs, and many used inconsistent field labels or handwritten notes.
The extraction system combined optical character recognition, computer vision, and natural language processing using Azure AI Document Intelligence to interpret document content. Text was first extracted from scanned and digital files. The system then analyzed layout and context to identify where patient identifiers were likely to appear. Instead of relying on fixed templates, the approach focused on understanding surrounding labels, positioning, and language patterns to locate the correct fields across different formats. This flexibility meant the system could adapt to new document types without requiring constant reconfiguration.
Improving accuracy without slowing operations
As the system matured, it demonstrated strong performance on referral documents commonly used by the company. On a representative test set, the automated extraction achieved over 90% accuracy for the required patient identifiers, allowing most referrals to be processed without manual data entry.
The system also handled documents it had not previously encountered. When tested on new and unseen document types, it maintained reliable performance, showing its ability to generalize beyond the original training set. For cases where information was missing or unclear, the system flagged records for human review, ensuring accuracy without slowing down overall throughput. By automating routine extraction tasks, intake teams were able to focus on exceptions and quality checks rather than repetitive data entry. This shift reduced errors caused by manual processing and improved consistency in patient lookup across systems. Processing time dropped by approximately 75%, and data entry errors decreased noticeably within the first month of implementation.
Ready to see results like these?
Tell us about your challenge and we'll show you how we can help.