Home Frontier Tech AI & ML The new bottleneck in enterprise AI sits inside the document

The new bottleneck in enterprise AI sits inside the document

- Advertisement -

As foundation models become more capable, accessible, and interchangeable across providers, many organisations are discovering that model performance is no longer the primary constraint on AI outcomes. Instead, the bottleneck has moved earlier in the pipeline, into the document layer that feeds AI systems in the first place.

IDC predicts that by 2027, 45% of the top 2,000 organisations in Asia-Pacific will adopt performance-intensive, software-driven, scale-out storage infrastructure and unified data management to accelerate insights for AI and analytics. This reflects the growing need to manage expanding datasets, unstructured data, and sovereignty requirements.

In conversations with CIOs and CDOs across financial services, healthcare, and telecommunications, the same observation comes up repeatedly. The challenge is not how the model reasons, but what the model is reasoning about.

- Advertisement -

The real problem: Unstructured data and document intelligence

Organisations in regulated financial services can encounter issues when manually gathering fragmented data from a wide array of documents. AI models can gather information from structured and unstructured sources, automating the data-extraction process. Using LLMs, organisations can then create tailored source-of-wealth write-ups while safeguarding customer information, before using an agentic AI framework to generate reports and validate the output. The final stage can include human validation to ensure accuracy and alignment with regulatory guidelines through control-function review.

In regulated enterprise environments, most critical data does not live in clean, structured warehouse tables. It lives in unstructured formats: PDFs, scanned filings, claims schedules, contracts, financial statements, lab reports, and rate exhibits. This is the data that feeds AI systems.

Document intelligence refers to the process of converting unstructured data into structured, usable inputs for AI models. When this step fails, the consequences ripple through the entire AI pipeline, ultimately resulting in the model generating a confident answer from a confidently wrong input.

At that point, even an advanced model cannot compensate for the underlying error. Cloudera’s “Data Readiness Index 2026” study found that 19% of APAC respondents cited data-quality issues as a main contributor to poor returns on AI initiatives, highlighting how the quality of the document pipeline can limit an AI system’s accuracy and directly affect its return on investment.

When parsing errors become business risks

The impact of poor document parsing shows up directly in business outcomes. The effects are concrete, measurable, and often underappreciated at the executive level. This is what makes document intelligence more than a technical challenge. For many organisations, it has become an operational and financial one.

Parsing errors are rarely visible when they occur. Instead, they compound downstream through workflows, decisions, and business processes. By the time the issue surfaces, the cost of remediation is often far greater than the cost of prevention.

  • In financial services, a single misread value in a fund administrator’s capital account statement can cascade into downstream errors in underwriting or reserving models. These errors can carry regulatory implications and lead to costly remediation efforts that often run into the millions.
  • Healthcare organisations continue to rely heavily on manual document abstraction for claims, remittance advice, and clinical documentation. This is driven in part by the structural complexity of the data, and in part by strict requirements around protected health information. Manual document abstraction is consistently one of the largest line items in health data operations budgets.
  • In telecommunications, vendor interconnect billing and service-level agreements often contain complex rate tables that few systems read accurately at scale. These parsing gaps can result in revenue leakage. According to “Revolutionizing telecom revenue assurance: the AWS AI-driven framework for next-generation solutions”, an AWS blog post published in June 2025, revenue leakage affects up to 10% of telecom revenue worldwide.

The trade-off: Accuracy vs. data sovereignty

In addition to accuracy, regulated enterprises must consider where AI processing takes place.

Over the past several years, most of the enterprise AI inference stack has steadily moved into controlled customer environments: virtual private clouds, on-premises infrastructure, or sovereign cloud regions. Models, vector stores, orchestration layers, and observability now operate within the same governance controls as the underlying data. Document parsing has been the forced exception.

Historically, the most accurate document-processing options were delivered via SaaS APIs only, which left regulated customers choosing between accuracy and sovereignty: sending sensitive documents to third-party tools for higher accuracy or keeping data within enterprise boundaries and accepting weaker performance.

A maturing category for enterprise document intelligence

Some document-intelligence systems can operate within customer-controlled environments, while benchmarks can be used to evaluate their performance on complex tables and highly structured business documents.

In this architecture, document intelligence operates within the same environment as the rest of the data pipeline, where structured and unstructured data are subject to the same security, lineage, observability, and governance controls from ingestion through inference.

Within this pipeline, document intelligence can parse unstructured documents, convert them into structured data, embed and retrieve relevant context, and generate outputs using AI models within the same controlled environment.

What this means for regulated industries

Document intelligence has several applications in regulated industries.

  • In financial services, applications include processing 10-K analysis and filings, fund administration records, bordereaux, claims schedules, and actuarial reports within governed environments, as well as using the resulting data in downstream agents and human review workflows.
  • In healthcare, applications include clinical trial data extraction, lab panel ingestion, and explanation of benefits processing within the same environment as structured protected health information.
  • In telecommunications, applications include interpreting interconnect agreements and billing structures to identify potential revenue leakage within complex rate tables.

The future: The sovereign AI stack

As models continue to converge in capability, organisations are paying greater attention to how they process and govern unstructured data throughout the data-to-inference pipeline.

The accuracy of document processing can affect the accuracy and ROI of AI systems. Regulated industries must also consider data-sovereignty requirements when determining where AI systems run.

- Advertisement -