Sep 13, 2026
Why capturing knowledge matters more than capturing data for AI
Your business captures its data and forgets how it decides.

Capturing knowledge matters more than capturing data for AI because a model can already read your data and can't read your knowledge: how your organisation works, what it's trying to do, who owns which decision, and why a number means one thing in finance and something else in sales. Almost none of it is written down anywhere a machine can reach, and it's the half that decides whether an answer is any use.
We've spent twenty years as an industry getting good at the first job and close to no time on the second, so most organisations now hold an enormous, well-governed account of what happened. Almost none of them hold an account of how they work. That ordering made sense while data was hard to move and expensive to keep. It's stopped making sense, and the reason is the thing everybody is currently buying.
So here's a question worth putting to your own team. If you connected a capable model to every system you own tomorrow, with full permissions and security waving it through, what would it still get wrong?
In our experience working with clients, the answer is most of what matters. You'd get a fluent account of last quarter and no idea which of your three definitions of an active customer you meant, or that one site's numbers have been off since a system migration and everyone who uses them knows to adjust, or that margin fell in one region because somebody made a pricing decision on purpose for a reason that still holds. None of that is in the warehouse. All of it is the difference between a useful answer and a confident one. Hold on to the migration one, because it comes back later.
What is the difference between data and knowledge in a business?
Data is what happened, and knowledge is how the business works and why it did what it did. Revenue by region, tickets by week and stock by depot are data. The reason you hold more stock at one depot, the assumption underneath the forecast that sets the level, and the fact that your commercial director overrides that forecast twice a year for reasons she has never written down, all of that is knowledge.

The first gets captured by default, because the systems running your business capture it as a side effect of running. The second gets captured only when somebody chooses to, which is to say hardly ever, and it leaves the building at the end of every meeting. Side by side, the two look like this.
Data | Knowledge | |
|---|---|---|
What it is | What happened | How the business works, and why it did what it did |
An example | Revenue by region, tickets by week, stock by depot | How "active customer" is defined, who owns that definition, and why one region is priced differently |
Where it lives | Rows and columns, in a warehouse | People's heads, documents, and the notes inside your CRM, your ERP and the email thread with a supplier |
How you capture it | By default, as a side effect of trading | Only when somebody decides to write it down |
What a model can read today | All of it | Almost none of it |
What it gives you | A description of the past | The context that makes a description mean something |
Why is your data worth less than it was?
Because everybody's tools can now read it, and a capability everybody has stops being a capability. Ali Ghodsi, the chief executive of Databricks, gave the industry's version of this to Forbes in August when he said there's "a major gap between the intelligence AI possesses and the impact it is having", and his diagnosis is that the missing piece is context, because a model can't reason about a business problem without reaching that business's records, its internal rules and its permissions. We think he's right, and we'd add that his fix covers about half the ground.
Records, rules and permissions are the part of context you can point an engineer at, and that work is real and it's being done well by people with a great deal of money behind them. The harder part is everything governing how the organisation makes up its mind, and that part was never entered into a system in the first place, so no amount of connecting improves it. You can hand a model every row you own and it'll still be reading your business from the outside.
Where does a company's knowledge actually live?
Your business knowledge lives in three places, and only one of them is searchable. Most pivotal insights live in people's heads, whereas some recorded decisions live in documents, and only some key details live in the notes inside your CRM or your ERP and the emails going back and forth between you and a supplier. That last one is knowledge sitting inside a system you pay a licence for every month, and it is still only partial knowledge that's often left invisible: most of your AI will not index it, the fields mean whatever the last person assumed, and nothing downstream reads it in a structured manner.

Two pieces of our own work make this better than the argument does. Both were advisory and delivery engagements, and neither of them was our platform, which you should know before reading them as product proof.
When we built an AI coaching service with ICE Creates for NHS patients in Wales, almost none of the raw material was data in any conventional sense. It was counselling transcripts sitting as audio files and Word documents in a SharePoint folder, and the number that mattered at the end was whether the thing had absorbed the expertise inside them: 73% of its coaching aligned to evidence-based clinical technique, which is published on our own case-study page. It stayed a knowledge problem all the way through.
On separate work with IM Group, a large part of the job was an enterprise data audit that went through more than five thousand custom data fields. Nothing gets built on five thousand fields until somebody has decided what each one means, which of them are dead, and which two are the identical thing under different names, and none of those decisions exist anywhere until a person writes them down. It's unglamorous and it's the actual work.
What is a semantic knowledge layer?
Semantic knowledge layer is a written account of how your organisation operates, held somewhere software can read it, with a history of what changed and who changed it: your definitions, your rules, your processes, who owns what, and the relationships between all of them. The word semantic is doing real work there. It means the layer holds meaning as well as structure, so it knows "active customer" is a decision your business made and it can tell you who made it.
Here's one entry, filled in and illustrative. It looks unremarkable, and that's the point.
Term: active customer
Definition: placed a chargeable order in the last 90 days. Excludes sample orders, internal accounts, and the two legacy reseller codes.
Owner: commercial director
Why it is defined this way: the 90-day window was set in 2023 to match the reorder cycle on the core range. It was never revisited when the subscription lines launched.
Where it disagrees: finance counts an account as active for twelve months after the last invoice, so the finance number is always higher. Board reporting uses the finance version.
Last reviewed: March, by the commercial director and the head of data.
We will not call a first build finished under sixty agreed definitions, because below that a system is still guessing at the terms when it answers a commercial question. The entry above is what we would want each of them to carry. It's a floor, and plenty of businesses will need considerably more.
The fourth and fifth lines are the ones doing the work here. A data catalogue records that the field exists; it doesn't record that the window was set for a reason which has since expired, or that two departments disagree and one of them wins in the board pack. Those are decisions the organisation made, and they're what a new hire spends six months absorbing, and what a machine is left to guess at.
We should be straight that context layers are crowded ground already. The largest data platforms are building them and saying so in public, and anybody who tells you they invented the idea of giving a model business context is a few years late to it. What we find thinner on the ground is a layer holding how the organisation decides as well as what it stores: the strategy it's working to, the processes and the workflows, the people, the culture. Writing that down is harder than mapping a schema, and it's where we think the value is. We'd put the split at about a third technology and two thirds getting people in a room to agree, though that's our experience of it rather than a measurement.
Why do models still get your business wrong when your data is connected?
Because a language model gives you the most plausible continuation of your question, and the likeliest answer is not always the true one. It has no way of telling the two apart. The limitation is in how these things work. Nobody is about to fix it, and we would sooner say so than sell you the tidier version where context turns out to be the cure.

It is also why guardrails and confidence matter. An answer that arrives with its definitions and its sources attached is one you can test. A confidence score on its own tells you the model is unsure; the evidence tells you why, and which of the two you get decides whether anyone catches the mistake.
Give a model a thin slice of your business and it can only understand a thin slice of your problem. The best prompt engineers get the best answers out of the same tools everybody else has because they are hand-feeding context that nobody ever wrote down. It works. It also does not scale, because you cannot sit a prompt engineer next to every decision your leadership team makes, and writing the context down properly, so the software reaches it every time, is months of work.
Most of the wrong answers we meet have nothing to do with invention. They are correct arithmetic on the wrong population. The model averaged over accounts you would have excluded, or used the finance definition when you meant the commercial one, or counted a month your business has always treated as an outlier.
Every one of those is right and useless. Every one of them would have been avoided if somebody had written the definitions down and handed them over first. Which is the site with the bad numbers from the top of this piece, and the second half of Ghodsi's diagnosis, the half we said his fix does not reach.
What are the rules for a business using AI to make commercial decisions?
In the UK there are almost none, and no law tells you how to decide.
Britain has stopped short of an AI Act. The working position is that existing regulators handle AI inside their own patch, so the FCA looks at financial services and the ICO looks at personal data, and nobody holds a rulebook for how your leadership team reaches a decision.
Europe did pass an AI Act, and then moved it. The rules for high-risk AI systems were meant to apply from 2 August. A few weeks before that date, the EU pushed them back to December 2027, and to 2028 for AI built into regulated products. The transparency rules landed on time, so you still have to tell people when they are talking to a machine. Everything else slipped by sixteen months, weeks before the deadline it was meant to meet.
Most businesses reading this were never in scope anyway. High risk means biometrics, critical infrastructure, education, employment, essential services, law enforcement, migration and justice. Forecasting demand or working out why margin moved is nowhere near it, and we would rather tell you that than let a compliance worry do our selling for us.
So for a good number of you, nobody is coming. The reason to write this down is in the frameworks themselves, not the law.
What do the governance frameworks ask you to document?
What these regimes ask for is worth reading even if they never apply to you. Europe's high-risk rules want to know how your data is governed and where it came from, how the system was built, logs of what it did, a plain statement of what it can and can't do handed to the people operating it, and a named human with the standing to override it. America's NIST framework, entirely voluntary, organises the same ground under Govern, Map, Measure and Manage. ISO 42001 asks you to record which controls you applied and which you did not. Three bodies, none of them obliged to agree with the others, and one shared instruction: write down how the thing works, who owns it, what it was told, and what it did. That is most of a knowledge layer, described by people trying to describe a safety regime.
Most of it, and the gap is where the money is. Every one of those regimes documents the system, and all of them stop at the system's edge. None of them covers whether "active customer" carries three definitions in your business, or which department wins when finance and commercial disagree. You can satisfy every logging requirement in full and hold a flawless record of a model answering the wrong question. Compliance buys you the audit trail and leaves the meaning exactly where it was.
Which is why the order matters more than the deadline. Build the documentation for the auditor and you get a record built to be filed, because the expensive ingredient, a written account of who owns which decision on what definition, was outside the ask. Build it for the business and you get something people use in a Tuesday meeting, with the audit trail falling out of it, because a system that can already show its definitions and the sources under them can already answer most of what any of them asks.
Sixteen extra months only helps the organisations that were doing this for a regulator in the first place. Everyone building it so the AI works is on the same timetable they were on before the deferral, because the thing making the AI wrong did not get an extension.
Does data stop mattering?
No, and the argument was never that it does. Data is the floor and you can't reason about anything without it. What's changed is only that data has stopped being the part that separates you from anybody else, because everyone's tools can reach it now, and knowledge is the scarce input because capturing it was never anybody's job. If your numbers don't reconcile today, fix that before you worry about a knowledge layer.
What have we built for Decision Intelligence?
This is the problem Unloq was built around, and the knowledge layer is most of what we mean by decision intelligence, so it's fair to ask where the edge of it is. We'd sooner draw that line ourselves than have a prospect find it in a demo.
What we offer today to our clients is the knowledge and context layer. We write it with you rather than bolt it on as a block of prompt text, and it keeps a history, so you can see what a definition said last quarter and who changed it. The distinction matters if you have been sold a chatbot before. Every answer comes back carrying its own working: the plan the system followed, the query it ran, the documents it retrieved with the page and the excerpt, how the numbers were added up, and whether the checks passed. Figures get checked against the query that produced them before anybody sees them, and in Phase 1 the system doesn't invent a number, so where it can't stand a figure up it tells you.
That version tells you what happened and why. It answers questions about your business with the evidence attached and it ranks the reasons a number moved. Ranked recommended actions, forecasts with confidence bands and approval routing are Phase 2, from October. External market signals are a Phase 2 module in build. And the piece we want most, a full record of every decision with the impact it was expected to have set against what happened, has a defined schema and no delivery date, so treat that one as our intent. We've written separately about why three quarters of UK AI adopters say it's working and almost none can prove it, which is this same gap seen from the far end, after the decision instead of before it.
The real limit, though, is upstream of all of it, because no product can install your organisation's knowledge for it. We can build the layer, keep it current and put it where software reads it, and somebody inside your business still has to say what an active customer is and which department wins when two definitions disagree. That's the tension we haven't solved and we're not certain anyone can, because the capture is a habit before it's a technology.
Where does capturing knowledge start?
With the terms, and with about an hour. Take the six numbers your leadership team argues about most, and for each one write the definition, the person who owns it, the reason it's defined that way, and where another department disagrees. Six is enough to be useful and small enough to finish in one go.
You'll learn two things quickly and one of them is uncomfortable. The first is that several of your definitions have no owner, which everybody suspected. The second is that at least one was set for a reason that stopped applying years ago and slipped past everybody, because the reason was never written down next to the number.
That document is the start of your knowledge layer and it's worth something whether or not you ever buy software from anybody. It's also, near enough, the first page of what a regulator would want from you one day, which is the second reason to write it and much the duller one. Mostly it decides whether AI works in your business, which is a strange fate for a page of definitions nobody wanted to own.
If you'd rather not start from a blank page, bring us a question your business argues about. We'll show you what answering it properly takes: the definitions it rests on, the evidence underneath it, and what the system does when it can't be sure.




