Simplified AI Model Selection with a Document AI Benchmarking Platform

Challenges
Too many AI models, with no clear way to choose the right one for each document and use case. Testing models individually was time consuming and costly, requiring separate APIs, development effort or GPU infrastructure. Vendor benchmarks failed to reflect real-world performance, making it difficult to compare accuracy, speed, cost and data-control requirements.
Outcome
Compared nine top AI models on accuracy, speed and cost in one platform. Enabled model testing without separate environments, across cloud, GPU-hosted and open-source options. Provided an evidence-based way to shortlist the right model for each document processing use case.
Solution
Document AI Benchmarking Platform
Challenges
Solution
Technology Stack
Outcomes
A UK-based retail giant deals with many types of documents across its business. Invoices, receipts, bank statements, product documents and ingredient information are just a few examples. The information inside these documents often needs to be extracted before it can be searched, analysed and used in applications such as RAG.
The challenge was not finding a document AI model. There were already plenty of them. The real struggle was finding which model the organisation should actually use for a particular document and use case. Different models performed differently. Some were more accurate. Some were faster. Some were cheaper. Others offered more control over where the data was processed.
Testing them individually was also difficult. They would need separate subscriptions, API access, development work or GPU infrastructure just to find out whether a model was suitable. That was time consuming and costly too.
Cloudaeon worked with the organisation to remove that guesswork. Instead of recommending one model upfront, we built a Document AI benchmarking platform, where users could compare, test and evaluate different document processing models from one place. The engagement changed the conversation from "Which AI model looks best?" to "Which model performs best on our documents and for our priorities?"
Challenges
Too many models, but no simple way to choose
The market already had several OCR, Document AI and multimodal models. That gave choice, but not clarity. A model that works well for an invoice may not be the right choice for a scanned document, form or product document. The best option can also change depending on what matters most to the user. For one workload, accuracy may be critical. For another, processing speed may matter more. At scale, cost per page or infrastructure cost can become equally important. The organisation needed a way to compare before committing to a model.
Testing each model separately required effort
Model evaluation was fragmented. To test a cloud service, a team needed API access and paid usage. To test an open-source OCR model, somebody needed to write code, create GPU infrastructure, deploy the model and then validate whether it worked. That created engineering work before they had even decided whether the technology was suitable. Moreover, the team then had to repeat the same exercise for the next model.
Vendor benchmarks did not answer the real question
Most models publish their own accuracy or benchmark results. Those scores are useful, but they are based on the vendor's own test datasets. That does not prove how the model will perform on an organisation's actual documents. The retailer therefore needed more than a global accuracy number. It needed a way to process its own documents.
Cost, speed and accuracy had to be considered together
A more expensive model is not automatically a better model. If two models produce the required level of extraction accuracy, but one costs significantly less to run, the cheaper model may be the better engineering choice. The same applies to speed. Some workloads need the most accurate extraction possible. Others need a fast response. Some need the lowest practical processing cost. The platform therefore had to support model selection based on the actual requirement, rather than assuming one model should be used everywhere.
Data handling also mattered
The organisation did not want sensitive documents leaving its own environment. That meant the solution needed to support different deployment choices. Users should be able to evaluate cloud API services where appropriate, while also having the option to work with open-source or GPU-hosted models where tighter control over data was required.
Root Cause Analysis
Cloudaeon did not start by selecting a preferred OCR or LLM and building around it. We first looked at what was making document AI selection difficult. The root problem was not document extraction itself. It was the lack of consistent decisions around those tools. There was no single place where a user could:
Understand what different models were designed for
Compare accuracy, speed and cost
Test models on their own documents
Evaluate extraction quality independently
This changed the direction of the solution. Instead of building another document extraction model, Cloudaeon built an interface for model discovery, testing, benchmarking and evaluation.
Solution
Cloudaeon developed a web-based Document AI benchmarking platform that brought different document processing models into one workflow.
How We Delivered
The step by step approach was followed:
Model discovery in one place: The platform provides users with a catalogue of available document processing models. Each model can be supported by a model card containing information such as:
Model specifications
Where the model performs well
Strengths
Limitations
Recommended usage practices
Source country
Deployment type
The platform groups services into different options, including free services, GPU-accelerated models hosted on Cloudaeon infrastructure and cloud services accessed through platforms such as Azure AI Foundry. This helps users understand the choice before processing a document.
Selection based on the actual requirement: Users can identify their document type and select what matters most to them. The available priorities include:
Best accuracy
Fastest processing
Lowest cost
Users can also combine priorities, such as accuracy and cost or accuracy and speed. The platform then narrows the available model choices accordingly.
AI-assisted model recommendation: For users who do not know which model to start with, the platform can suggest suitable models. It can recommend, for example, a model that offers the best overall fit or one that provides a lower-cost option. The user still controls the final choice.
Document processing without building each integration first: Once a model is selected, the document can be processed directly from the interface. Users can upload a document and either extract the document content generally or provide an instruction or prompt to retrieve specific information. Some models extract the complete text first. That output can later support RAG or other information retrieval workflows. Instruction-based vision models can also interpret the document and return the information requested by the user. The platform was also designed to support multiple documents and batch processing.
Benchmarking based on actual testing: A key part of the solution is that the scores shown in the platform are not simply copied from model vendors. Cloudaeon tested models using available documents and used those results to provide comparative scoring. The platform also includes a benchmark view that helps users understand how different models compare.
Independent evaluation of extraction quality: The evaluation capability goes further than displaying an accuracy percentage. A processed document can be loaded alongside the expected text or ground-truth file. The platform can then compare the extracted output with the expected result. A code-based evaluation approach, OmniDocBench strategy, can evaluate extraction without relying entirely on another LLM to judge the answer. This is important because using an LLM to judge another model introduces another layer of uncertainty.
Flexible infrastructure and data control: The platform supports different operating models. Cloud services can be accessed through APIs. GPU-based models can run on Cloudaeon infrastructure. Open-source models can also create a path for organisations that want tighter control over where documents are processed. If a customer later selects a particular model for production, Cloudaeon can help establish the required infrastructure, including GPU-based deployment where appropriate. The benchmarking platform therefore helps teams validate the technology before making a larger infrastructure or service commitment.
Technology Stack
Azure AI Foundry
Azure Document Intelligence
Surya OCR
Mistral
GPT 5.1
MinerU
PaddleOCR-VL
Qwen-VL
Granite Docling
Databricks document extraction/model services
Python-based document processing models
Vision-language and instruction-based document models
Cloudaeon-hosted GPU infrastructure
Azure Kubernetes Service (AKS)
API-based cloud model integrations
OmniDocBench-based code evaluation
LLM-based evaluation methods
RAG-ready text extraction workflows
Outcome
The organisation gained a single platform to compare the top nine AI models based on accuracy, speed and cost without relying only on vendor benchmarks.
Teams could test nine different models without building separate environments for each one, while retaining flexibility across cloud, GPU-hosted and open-source options.
The organisation could leverage an evidence-based method to shortlist the right model for each document processing use case.
Conclusion
With dozens of OCR, vision and document intelligence models available, choosing one based on vendor claims or reputation can quickly become expensive. The right model depends on the documents being processed, the accuracy required, the expected speed, the cost profile and where the data is allowed to go. Cloudaeon's work with the retailer created a more practical way to make that decision.
If your organisation is evaluating Document AI, OCR or multimodal models, Cloudaeon can help you benchmark the options before you commit to a production architecture. Talk to an expert about evaluating the right Document AI approach for your documents and workloads.
