All articles Insights · Data

Getting Your Data Ready for AI: A Practical Checklist

Neatly cabled server racks in a data centre

When an AI project stalls, the model is rarely the problem. The data is. It is scattered across systems, nobody is sure which version is right, and the one field the model needs most was filled in differently by every team. The good news is that getting data ready for AI is mostly organised, unglamorous work, and you can start before you have chosen a single model.

The checklist

1. Can you get to it?

List every source the use case needs: ERP, CRM, spreadsheets, sensor feeds, documents. For each, note how you can read it (API, database, export) and how often it changes. If the only way to get the data is a monthly spreadsheet emailed by one person, fix that first.

2. Do you know what it means?

Write down what the important fields mean, in plain language, with an owner for each. “Order date” can mean the date the customer clicked buy, the date it was confirmed or the date it shipped. Models do not know the difference unless you tell them.

3. Is it good enough for this decision?

Check completeness, duplicates, obvious errors and how consistent formats are across sources. You do not need perfect data. You need data that is good enough for the decision you want to automate, and you need to know where the gaps are.

4. Do you have the labels?

Supervised models learn from examples of the right answer: which invoices were fraudulent, which machines failed, which tickets were urgent. If those answers are not recorded, plan how to capture them. Sometimes the first project is simply adding a field to a form.

5. Is it safe to use?

Identify personal and sensitive data early. Decide what the model is allowed to see, where the data may be stored and processed, and how long it is kept. Under GDPR and similar laws this is not optional, and it is far easier to design in than to bolt on later.

6. Can it flow reliably?

A model trained once on a clean extract is a demo. A model in production needs a pipeline that delivers fresh, checked data every day, alerts someone when a source breaks and keeps a history so you can retrain and audit.

Do you need a data platform?

For a first use case, a well-built pipeline into a single database is often enough. Once several AI and analytics projects need the same data, a lakehouse pays off: one governed place where raw and cleaned data live together, with access rules you control. We usually let the second or third use case justify the platform, rather than building it in advance.

Integration is half the job

Much of the effort in data readiness is moving and mapping data between systems that were never designed to talk to each other. We built our own integration tool for exactly this reason, because the same mapping problems appear on almost every project.

A sensible order of work

Done this way, data readiness stops being a giant programme that never ends and becomes a series of small, useful steps, each one paid for by a working AI feature.

Frequently asked questions

Do we need a data lake before starting an AI project?

Not necessarily. Many first projects run on a well-prepared extract of the data they need. A lakehouse becomes worth it once several AI and analytics projects need the same data.

How clean does our data need to be?

Clean enough for the decision you want to make. Fix the fields the model depends on, document the known gaps and measure how quality affects results, rather than trying to perfect everything first.

Who should own data for AI?

Each important dataset needs a named business owner who understands what the fields mean, plus an engineering owner who keeps the pipeline healthy.

Planning something similar? Talk to our engineers or see our AI development services.

Keep reading

More from our engineers.

View all blogs