How We Built a Smart Medical Document Processing System and Improved Accuracy & Efficiency

Introduction

The client came to us with a simple goal: make it easier for people to manage their medical documents.

The idea was to build a platform where users could securely upload their medical records, keep everything organised, and access their documents whenever they needed them through a mobile app. An admin portal would provide the tools needed to manage the platform and its data.

But the requirement went beyond simply storing files.

The system needed to read and understand uploaded medical documents, extract their content, support multiple languages, and classify documents intelligently. It also needed to do all of this without depending on direct integrations with hospitals or healthcare providers.

That is where the real challenge began.

The Challenge

At first, the requirement sounded fairly straightforward: upload a document, extract the text, and make it usable.

In practice, medical documents are rarely that simple.

PDFs could contain multiple pages, complex tables, different layouts, and content in languages such as Spanish and German. We needed to preserve as much of the original structure as possible while still making the extracted information accurate and useful.

One of the biggest challenges was processing PDFs. We had to split each PDF into individual pages, convert those pages into images, and then process them for OCR, validation, and formatting.

During testing, we came across several issues:

  • Tables were not always extracted correctly
  • Some multilingual content was interpreted inconsistently
  • Certain words were missed or incorrectly recognised
  • Some table values were returned as empty
  • Complex document layouts were difficult to preserve

There was another problem we had to solve as well.

The platform needed to identify whether an uploaded image was actually a medical document. Using an LLM for every validation request would have worked, but it would also increase processing time and token usage, making the solution more expensive to operate.

We needed a better approach.

Our Approach

Instead of trying to solve everything with a single AI model, we broke the problem into smaller parts.

We looked at where traditional OCR and image-processing techniques would work well, and where an LLM could actually add value.

This led us to a hybrid approach.

  • Use OCR and image-processing tools for extracting the actual document content
  • Use the LLM where intelligent formatting or interpretation was required
  • Use rule-based and keyword-based checks for basic medical document validation
  • Keep the LLM out of parts of the workflow where it was not necessary

This helped us focus the AI where it provided the most value instead of using it for every step.

How We Built It

Phase 1 – Understanding the Problem

Before getting into implementation, we identified the areas most likely to cause problems later.

These included:

  • Multi-page PDF processing
  • Maintaining document structure and layout
  • Supporting multiple languages
  • Extracting information from tables
  • Validating medical documents
  • Controlling LLM usage and processing costs

We then designed the processing pipeline so that OCR, validation, and formatting were handled as separate stages.

This separation made it easier to test each part independently and identify where improvements were needed.

Phase 2 – Building the Processing Pipeline

We implemented the backend using Node.js, with Google Vision API handling OCR and image analysis and OpenAI supporting intelligent formatting and validation.

The processing flow was designed around a few key principles:

  • PDF pages are converted into images before processing
  • OCR handles the initial text extraction
  • Additional processing improves results for multilingual documents
  • Google Vision and rule-based checks help determine whether a document is relevant
  • The LLM is used selectively for tasks that require deeper interpretation or formatting

This allowed us to reduce our dependency on the LLM while still benefiting from its capabilities where they mattered most.

Phase 3: Testing and Improving

A large part of the project involved testing different document types and finding the cases where the system struggled.

We tested OCR across multiple languages and document formats, reviewed the output, and then refined the processing logic.

We also focused on table extraction because this was one of the areas where standard OCR alone was not always enough.

At the same time, we monitored token usage and looked for opportunities to remove unnecessary LLM calls.

The goal was not simply to make the system work. It was to make the system work reliably without making every document unnecessarily expensive to process.

What Changed

Through multiple rounds of testing and optimisation, the solution moved well beyond the original OCR requirement.

We achieved:

  • Better document extraction across multiple languages
  • More reliable identification and validation of medical documents
  • Improved handling of complex document structures
  • Reduced unnecessary LLM usage
  • Lower processing overhead
  • A more structured and maintainable document-processing pipeline

The biggest improvement came from treating OCR, validation, and AI interpretation as different problems rather than trying to solve everything through one technology.

Visual Proof

The development process can be demonstrated through:

  • Before-and-after document extraction examples
  • OCR accuracy comparisons across languages
  • Document validation workflow diagrams
  • Processing and performance comparisons
  • Examples of complex tables and how their extraction improved

These visuals help show the difference between the initial output and the results achieved after optimisation.

What We Learned

One of the biggest lessons from this project was that AI does not always need to be involved in every step.

For some tasks, traditional image processing and rule-based logic were faster, cheaper, and more predictable. For others, the LLM provided the flexibility we needed.

A few key takeaways stood out:

01

Combining AI with traditional techniques can produce more reliable results than depending on a single technology

02

Multilingual OCR needs continuous testing and optimisation

03

Reducing unnecessary AI calls can have a significant impact on operating costs

04

Complex documents need to be tested against real-world edge cases, not just standard examples

05

Iterative testing is essential when working with document-processing systems

If we were starting the project again, we would spend more time upfront evaluating advanced table extraction approaches. That would have helped reduce some of the rework we encountered during later testing.

The Outcome

What started as a requirement to extract text from medical documents became a much broader document-processing solution.

The final system combines OCR, image analysis, rule-based validation, and AI to create a more reliable and cost-efficient workflow for handling medical records.

More importantly, the architecture gives us room to continue improving the platform as we encounter more complex documents, languages, and use cases.

The project reinforced an important principle for us: the best AI solution is not necessarily the one that uses the most AI. It is the one that uses the right technology at the right point in the process.

bhavin parmar

Author

Faiz Syed

Faiz Syed is an ISTQB® Certified Test Engineer and Project Coordinator with hands-on experience in quality assurance across IT projects. He specializes in software testing, test case design, and quality processes, ensuring reliable and high-performing deliverables. Known for his collaborative and responsive approach, Faiz works closely with cross-functional teams to support smooth project execution.

Begin Your Next Phase of Digital Growth

Collaborate with a technology partner that understands your business goals and delivers solutions built for scale, security, and long-term success.

Schedule a Consultation