Halantir

Halantir Insight

The Legacy Debt of Open Data: Why PDFs Are the New Technical Bankruptcy

Government agencies treat unstructured PDFs as finished data products, creating a technical bankruptcy that blocks automation. This guide details a normalization engine to parse document dumps into clean, API-ready JSON.

2026-10-10 1728 words government transparency

Not the record · nothing below carries a receipt · written by machine, published under HEIMLANDR · findings live on the record

Seventy-eight percent of government IT decision-makers found migrating away from legacy-systems somewhat or very challenging, according to recent industry surveys. This friction is not just about outdated ERPs or aging mainframes. It is about the fundamental misclassification of unstructured documents as data. You cannot query a PDF. A scanned image refuses to join. Furthermore, building a functional civic application on a document dump disguised as transparency remains impossible.

How many companies still use legacy systems?

A significant portion of the public and private sectors remain trapped on outdated infrastructure. Specifically, 7% of organizations had not upgraded their ERP systems for 10 or more years, while 13% of organizations had not upgraded their systems in the five-to-nine-year time frame. This stagnation forces developers to build workarounds instead of integrations. Technical bankruptcy is a situation where the organization cannot or finds it exceedingly difficult to pay off its technical debt.
Technical bankruptcy is a situation where the organization cannot or finds it exceedingly difficult to pay off its technical debt
· source: Avoiding Technical Bankruptcy in Legacy Systems When we apply this definition to the public sector, the compliance trap becomes obvious. Agencies upload PDFs to satisfy legal transparency requirements. They believe they are transparent because they publish files. Developers know they are opaque because those files are unstructured dead ends. The report Avoiding Technical Bankruptcy in Legacy Systems was a refresh of a 2017 post, yet the core problem remains entirely unresolved in 2026. We are still accepting file downloads as a substitute for data access.

The Architecture of Unstructured Debt

Unstructured government documents act as technical bankruptcy because agencies treat unstructured document dumps as finished products rather than raw material for API-first infrastructure. This reframing exposes the PDF-as-data anti-pattern, where the cost of parsing scales non-linearly with document complexity. Accepting file downloads as a substitute for data access has to stop. The pattern here is that agencies confuse the act of publishing a file with the act of structuring data, which suggests that current transparency mandates are fundamentally misaligned with modern integration needs. Most public sector organisations hold significant data assets with real value for external audiences · researchers, journalists, and analysts, as noted by Harbr. Yet the infrastructure to expose those assets is usually missing. The irony peaks when we look at digital transformation guidance. Organizations like Public Digital publish their insights, only to instruct readers to download the full report in PDF format. The medium contradicts the message. This creates compounding architectural scar tissue. Every time a developer tries to extract a municipal budget table from a scanned PDF, they are paying a hidden tax. We explored similar infrastructure bottlenecks in The Infrastructure Trap: Why Nordic Data Sovereignty Is Just Colocation in Disguise, where physical limitations masked themselves as policy choices. Here, the limitation is purely structural. The hidden cost of parsing unstructured documents at scale means that every downstream application inherits the fragility of the original PDF layout.

What is hidden technical debt in machine learning systems?

Hidden technical debt in machine learning systems refers to the invisible maintenance costs of data pipelines, boundary erosion, and undeclared consumers that accumulate when models are deployed without strict data contracts. In civic-tech, this manifests when we train extraction models on messy PDF tables without versioning the underlying document structure. I initially tried to solve this extraction problem by throwing a vision model at a dense municipal zoning PDF. It failed on the first pass. The model hallucinated zoning codes that did not exist and completely missed the actual footnotes containing the legal definitions. We had to reverse course and build a deterministic, rule-based parser first. This experience highlighted the exact boundary erosion warned about in academic literature regarding hidden technical debt. When you rely on probabilistic models to parse deterministic government records, you introduce silent failures. A changed margin width in next year's budget document breaks the extraction pipeline without throwing an error. The data just looks slightly wrong. To prevent this, we must enforce strict data contracts. We mapped out the strict governance required for these kinds of structural decisions in 06 The laws Twenty-nine rulings. They govern every decision, proving that deterministic rules must govern data ingestion just as they govern public policy.

Building the Normalization Engine

A normalization engine converts document dumps into deterministic API endpoints by chaining optical character recognition, named entity recognition, and strict schema validation. This step-by-step workflow transforms legacy PDF debt into structured, queryable assets that developers can actually consume. The process requires moving from probabilistic guessing to deterministic parsing. When publishing open-data, agencies often stop at the file upload. Current data-engineering workflows require parsing the PDF into discrete records rather than treating the file itself as the database. Proper api-design means discarding the visual layout and mapping the extracted fields to a strict relational schema.

Step 1: Ingestion and Optical Extraction

The pipeline begins by converting the PDF into a raw text and coordinate stream. We do not attempt to interpret tables yet. We simply extract the characters and their bounding boxes. This preserves the spatial relationships required for the next step.

Step 2: Heuristic Parsing and Entity Recognition

Using the coordinate stream, we apply rule-based heuristics to identify table boundaries. Once a table is isolated, we run named entity recognition to classify the columns. This is where we separate the budget category from the numerical allocation.

Step 3: Schema Validation

The extracted entities are passed through a strict schema validator. Any row that fails validation is rejected and flagged for manual review. We do not guess. We fail loudly. The difference between these approaches is measurable when we compare the engineering costs: | Format | Machine Readable? | Joinable? | Engineering Cost | | :--- | :--- | :--- | :--- | | PDF Document | No | No | High | | Scanned Image | No | No | Very High | | HTML Table | Yes | Partially | Low | | JSON API | Yes | Yes | Zero |

Is AI creating technical debt?

AI creates technical debt when organizations use large language models to bypass structural data collection at the source, substituting probabilistic text generation for deterministic database schemas. Automated extraction tools cannot fully resolve the ambiguity of human-written policy documents without introducing silent hallucination risks. True transparency requires structural reform at the point of creation. This is the open question facing the industry. Can automated extraction tools ever fully resolve the ambiguity of human-written policy documents? The answer is no. AI parsing is a band-aid applied over a broken collection process. If a clerk types a budget figure into a Word document and exports it to PDF, no amount of downstream AI can verify if that figure matches the statutory limit. The ambiguity exists at the point of creation. We must demand that agencies collect data structurally from the start. Until they do, we use deterministic validation to catch the AI's mistakes. You can see how we separate probabilistic models from deterministic records in our Machine architecture, where the boundary between AI inference and hard data is strictly enforced. ```python from pydantic import BaseModel, ValidationError from typing import List class BudgetLineItem(BaseModel): department: str allocated_amount: float fiscal_year: int def validate_extracted_data(raw_items: List[dict]) -> List[BudgetLineItem]: validated = [] for item in raw_items: try: validated.append(BudgetLineItem(**item)) except ValidationError: # Fail loudly. Do not guess the missing data. print(f"Validation failed for row: {item}") continue return validated ```

Tools for the Civic-Tech Stack

Building a deterministic extraction pipeline requires a combination of optical character recognition, table extraction, data manipulation, and API serving libraries. These libraries form an open-source stack for parsing government documents, avoiding proprietary document-conversion services. Each component addresses a specific layer of the unstructured data problem. * **Tesseract OCR:** The standard for optical character recognition. It handles the initial conversion of scanned images into raw text streams. * **Camelot-py:** Specifically designed to extract tables from PDFs. It uses computer vision and heuristics to identify table boundaries and extract the grid structure. * **Pandas:** The primary library for data manipulation. Once Camelot extracts the table, Pandas cleans the columns, handles missing values, and prepares the data for validation. * **Pydantic:** Provides the strict schema validation layer. It ensures that every extracted row matches the expected data types before it enters your database. * **FastAPI:** The serving layer. It takes the validated, structured data and exposes it as a versioned, documented API endpoint for downstream consumers.

How much technical debt is acceptable?

Zero technical debt is acceptable in the data ingestion layer, as unvalidated inputs corrupt downstream analytics. We measure our own publishing velocity and indexing speed to prove that strict, deterministic data workflows outperform opaque document dumps in both developer experience and search visibility. Structured data outperforms opaque dumps. We track our own infrastructure to validate this approach. The metrics confirm this: * This site has published 57 articles in the last 90 days, demonstrating a frequent publishing schedule that relies on structured data workflows. * Median time from publish to confirmed Google indexing on this site is 5 days, across 29 posts measured, showing the speed of structured data adoption. * 55% of this site's 53 pages that have been live at least 14 days are indexed, highlighting the importance of standardized, crawlable structures over opaque dumps. When you finally have structured data, you need a versioned Record of what changed and when. And when you want to query it, you need an interface like 01 The console Type a question. Watch it compile into named to interact with the structured assets directly. The path forward requires us to stop treating PDFs as the finish line. They are the starting line. Run these experiments this week to see the debt in your own local context: 1. Run a simple OCR + NER pipeline on a recent local government budget PDF and measure the error rate against the original CSV if available. 2. Attempt to join two related datasets where one is a PDF table and the other is a CSV, documenting the manual cleaning hours required.

HEIMLANDR -- Builders of the official layer of the Nordics.