A model can summarize, extract, or respond only based on the information it receives. In companies, this information is distributed among shared folders, business applications, databases, emails, intranets, and PDF documents. They have different versions, rights, and owners. The AI project then reveals data problems that already existed.
Preparation does not consist of cleaning the entire company before starting. It consists of building, for a defined use case, a known, authorized, measurable, and maintainable data chain.
Start from the use case
A document assistant, an invoice extractor, and an agent who modifies orders do not have the same needs. It is necessary to specify:
- task;
- users;
- necessary sources;
- freshness;
- minimum quality;
- forbidden data;
- decision or action;
- expected proof.
This precision limits the inventory and avoids an abstract data platform without results.
Build a usable inventory
For each source, collect:
- name and system;
- business owner;
- technical manager;
- type and volume;
- format;
- update frequency;
- sensitivity level;
- authorized population;
- known quality;
- access method;
- retention rule;
- reference status.
The inventory must be maintained. A sheet frozen at the beginning of the project does not guarantee that the source used six months later remains legitimate.
The ten dimensions of preparation
1. Inventory
The necessary sources are identified and their scope understood.
2. Property
A person or function can validate the meaning, quality, access, and lifespan.
3. Access
The extraction is authenticated, documented, and sustainable. It is not based on a manual copy or a personal account.
4. Quality
The essential fields are complete, consistent, and representative. Defects are measured against the task.
5. Structure
Headings, sections, tables, relationships, and identifiers can be preserved or reconstructed.
6. Metadata
Language, date, version, status, type, author, entity, and rights accompany the content.
7. Version and freshness
The system knows which version is in effect, when it changed, and when the AI needs to be updated.
8. Rights and Privacy
Permissions can be reproduced in the AI chain and revoked.
9. Traceability
An output can be linked to the source, its transformation, and the version of the components.
10. Life cycle
The correction, archiving, and deletion propagate to indexes, caches, and logs.
Traceability of data from its source to its use by AI and its propagated deletion.
Task-related quality
A piece of data is not 'good' in absolute terms. To retrieve a procedure, the title, date, and status are essential. To extract an amount, the quality of the scan, the currency, and the table structure are dominant.
The dimensions can include completeness, accuracy, consistency, uniqueness, freshness, and representativeness. Each defect is related to its impact on the use case.
An annotated sample makes it possible to measure before processing the entire corpus.
Distinguish reference, copy, and draft
Several versions of a document may circulate. The pipeline must know which one is authoritative, which copies are archived, and whether a draft can be used. The date alone is not enough: a recent document may not be approved.
The priority rules are coded and verifiable. In case of conflict, the system signals rather than silently choosing.
Preserve the structure
Converting a PDF into plain text can mix columns, footnotes, and headers. Tables lose their relationships and titles lose their hierarchy. The extraction must produce a structured format with measured quality.
For complex documents, keep:
- page and contact details;
- titles and levels;
- cells and headers;
- lists;
- captions;
- links;
- file identifier.
This structure improves research, citations, and human oversight.
Metadata and context
Metadata allows filtering and explanation: entity, product, country, language, date, version, status, category, and rights. They must come from a reliable source or an extraction whose confidence level is maintained.
A model-generated classification does not have the same status as a field validated by the business. The source of the enrichment is recorded.
Reproduce the rights
Copying documents into a single index without permissions creates a potential leak. The authorization model must be understood: user, group, company, folder, site, role, or attribute.
Rights can be materialized in the metadata of indexes, partitions, or a query check. Changes and deletions propagate within a defined timeframe.
Technical accounts have minimum access and extractions are logged.
Minimize the data
A project should not ingest an entire file 'just in case.' Selecting the necessary fields and documents reduces risk, cost, and noise. Personal data or secrets can be masked, pseudonymized, or excluded.
Minimization also concerns prompts, logs, feedback, and evaluation sets. Development environments use synthetic or protected data.
End-to-end traceability
Each fragment or record must be able to be linked to:
- source and version;
- extraction date;
- transformation;
- tool and version;
- enrichment rules;
- index;
- outputs that cited it when necessary.
This chain allows investigating an error and reconstructing after an update.
Manage deletions
Deleting in the source is not enough. Copies, vector indexes, caches, temporary files, and backups must be deleted or invalidated according to policy. Historical outputs and logs follow a separate rule defined with the responsible parties.
The deletion pipeline is tested like the ingestion one. A stable identifier facilitates propagation.
Measure the freshness
The need varies: a policy can tolerate one day, a stock a few minutes. Freshness is tracked by source and visible to the user when it influences the response.
Periodic events or extractions must handle failures. An alert signals a delay rather than silently continuing with old data.
Prepare a reference game
Before AI, select representative examples and document the expected result. For extraction: fields and values. For RAG: questions and sources. For the agent: states and allowed actions.
This game is used to compare pipelines, models, and corrections. It includes difficult, incomplete, and forbidden cases.
Industrialize ingestion
The pipeline is:
- idempotent;
- versioned;
- observable;
- resumable;
- tested;
- secure;
- able to handle errors in isolation.
Format or source changes trigger an alert. A dashboard shows volume, errors, delay, duplicates, and coverage.
Light but real governance
Useful governance does not require a committee for every file. It defines owners, classes, controls, and an exception procedure. Important decisions are traceable.
The data catalog can start with the scope of the project and expand with the uses. It must remain connected to the systems, not become an obsolete parallel documentation.
Build a preparation backlog
Actions are often divided into three horizons:
30 days
Inventory, access, sample, classification and initial tests.
90 days
Reproducible pipeline, metadata, rights, measured quality, and deletion.
Foundations
References, governance, automation, observability, and sharing between use cases.
The pilot can start when the risks are controlled on a limited corpus, without waiting for overall perfection.
Data as a product of trust
A reliable AI depends on data whose meaning, owner, rights, and version are known. This preparation also benefits research, analytics, and standard integrations.
Partitech can audit the sources, design the pipelines, structure the metadata, and integrate the rights up to the RAG or the agent. The goal is a traceable and maintainable chain, not a massive copy of documents into a new index.
Let's talk about your project
Assess the readiness of your data for an AI project with Partitech. Contact Partitech.