Skip to the text

Off the Books The Share of Judgment

Essay I

The Share of Judgment

On what artificial intelligence reveals about expert organisations


Karim Louedec · July 2026 · ~32 min

Off the Books · Opus I · Essay I — datatropy.ai

Essay I

The Share of Judgment

On what artificial intelligence reveals about expert organisations

I. The Scene

A Tuesday morning, in the meeting room of a research centre. On the screen, an eight-page document produced in forty seconds by a machine. A synthesis. Twenty-three references, a line of reasoning, a recommendation.

Around the table, six people. Five are waiting. The sixth is the one who counts: thirty years in the job, a reputation that precedes her by three corridors, the kind of person of whom it is said that if she says so, it is so.

She reads. She sets down her glasses. She says: “No.”

Silence. Someone asks why.

She takes four seconds. She answers: “That is not how we work.”

There is no answer in that, a verdict without grounds. The meeting stops all the same, the project slips three months, no one objects. No one in that room has the slightest instrument for telling her she is wrong. Nor for telling her she is right, for that matter.

She was probably right, she often is. The organisation has therefore just taken a good decision for a bad reason. It did not judge the document, it followed a person. That the person is reliable changes nothing about what happened. The system works as long as she is there, and it leaves nothing behind it.

This scene shows what is found everywhere else. A whole organisation discovers, without putting it into words, that it has never been able to explain why a piece of work is good.

It knew something else. It knew who was good.

Between the two, an abyss.

The question being asked everywhere, is this AI reliable, is not the right one. The right question is: would you be able to say?

II. Displaced Evaluation

The system deserves its due before it is attacked.

In a laboratory, an R&D centre, a scientific division, a clinical research department, the evaluation of work has never been absent. It was displaced.

Displaced in time. A person is evaluated on entry, massively, once. Competitive examinations, a doctorate, publications, interviews, sometimes several days of selection. The investment is heavy and it is deliberate. Then it stops. The rest of the career runs on credit, credit that grows with seniority and is almost never revised downwards. It goes by preference to those who already have some.

Displaced in space. The real evaluation happens outside, through peer review, publication, citation. The organisation very rarely evaluates its own output. It waits for the academic world to do it in its place, with several months’ delay and a coverage that, in an industrial R&D centre, rarely exceeds ten per cent of what is produced there. The rest comes before no judge. Internal notes, decision files, watch summaries, portfolio arbitrations. Almost never.

Displaced in form, finally. What is left of internal evaluation belongs to social ritual rather than to criteria. The Thursday seminar. The project review. The scientific committee. Nothing is scored. There is discussion, there is nuance, and at the end someone whose voice carries more than the others steers the room.

Displaced towards the regulator, in some of these houses. An authorisation file is examined by appointed rapporteurs, who write a reasoned objection on the choice of comparator, on the population retained, on the primary endpoint and on the use made of the literature. That is judgment, written and signed, to which one must respond within a deadline. Thirty years of that regime have produced no internal evaluation capacity in the same houses. They learned there to manufacture the file the judge expects, and the judge stayed outside.

I will be told that these houses are covered in quality arrangements. That is true, and it is sometimes even an understatement. Protocols, standard operating procedures, sample traceability, statistical validation, regulatory review, cross-audits. Some industries are so saturated with them that control takes more time than the work controlled. They know perfectly well how to verify that a procedure has been followed.

Verifying that a procedure has been followed does not amount to evaluating a judgment. A protocol impeccably executed on a bad hypothesis produces an impeccable and useless result. Every box is ticked, every control passed, and the whole is worth nothing because the opening question was badly put. It is the hypothesis that has never been scored. The choice of question, the selection of sources, the construction of the reasoning, the decision to stop here rather than there.

My list is perhaps too convenient. Milestones with passage criteria written before a project is launched, design reviews, technology readiness levels, committees that rule each year on each family of projects with a written ground and a named decision-maker. Those are arrangements that decide on something other than a procedure followed. An end-of-phase milestone rules on the programme’s data and on its remaining cost. What it rarely does is score the quality of the reasoning of the person who proposed the programme. The file passes or it does not. No one writes that the opening hypothesis was badly built, and no one reopens the rating three years later to look at who was wrong and about what.

These organisations do evaluate: people, once; procedures, always; judgments, far too rarely.

It was not absurd. It was economic. Building a function that evaluates expert judgment costs money, takes time, and runs head-on into people one would rather not run into. As long as output stayed slow and rare, going without it was the right calculation. The entry filter did the work. When it takes fifteen years to train someone capable of producing a defensible hypothesis, and three weeks for that person to produce it, the scarcity of the producer more or less guarantees the quality of the product. One does not control what is rare and expensive to make. One controls the maker, once, on entry. There is just one consequence no one saw coming.

An organisation that does not evaluate its work does not develop the vocabulary to talk about it. After thirty years, the tool is missing and so are the words. Ask a scientific committee what separates a good literature review from a mediocre one. You will get adjectives. Rigorous. Solid. Serious. Well conducted. Complete. Honest.

These are appreciations, not criteria. Nothing can be done with them, no one can be trained, no disagreement settled, and certainly no machine scored.

Test it, it is instructive. Take two experts in the same field, show them the same document, ask them separately what they think of it. They will often converge on the overall verdict. Then ask them to write down the three criteria that produced that verdict. You will get two lists that overlap little, and often each will discover, listening to themselves, that they cannot say what they do. Ill will has nothing to do with it. They were never asked, and knowledge that is not stated does not become structured.

That is the hole. It is old, it is comfortable, and it stayed perfectly invisible until a Tuesday morning.

III. The Bottleneck Was Elsewhere

Generative AI was sold as a productivity gain. If my experts produce three times faster, I produce three times more. It is an industrial reading. It held for the production line, the back office, customer service, the call centre, and it was transposed without examination to a context where it does not hold.

The reason is structural. In an expert organisation, output has never been the bottleneck.

The bottleneck was selection. The capacity to decide, among all possible directions, which one is worth six months, a budget, a team. That decision rested on a small number of people, and it was rare because they were rare.

Multiply output by a hundred, and you have not multiplied performance by a hundred. You have moved the whole of the pressure onto the point that was never instrumented.

Generative AI multiplied what was not scarce and left untouched what was.

Take a literature review. Before, a researcher spent two weeks on it, drew out twelve references they had read, and their judgment on what counts was built into the text, invisible and beyond challenge. After, a machine produces thirty pages and a hundred references in an afternoon, of which ninety-five are real, three approximate, and two do not exist.

AI writes very well, the question is elsewhere: can you spot the two references that do not exist? If yes, you have just gained a factor of ten on raw output. A serious check will take a good part of it back, and the final count runs nearer three or four, which remains considerable. If no, you have just introduced into your internal corpus a slow poison, in a form that inspires more confidence than what it replaces. It is well written, it is well structured, it looks a great deal like serious work.

Take a second case, less obvious but possibly more revealing. A portfolio committee has to arbitrate between twelve research programmes. Before AI, each programme arrived with a fifteen-page file, produced in three weeks by its sponsor, and the committee spent half its session asking for details that were missing. Today, each file runs to forty pages, it is complete, it anticipates the questions, it is structured impeccably.

Is the committee better armed to decide? No. It is less well armed.

In the old world, what distinguished a good programme from a bad one lay as much in the quality of its argument as in its content, and that quality was an indirect signal of the quality of its sponsor. A confused file signalled confused thinking. The signal was coarse, unfair at times, and it carried information. Every file is now equally well written. The committee has to judge on substance, on substance alone, that is, to do explicitly what it has never had to do.

This is why the problem never appears where it is expected. It is looked for on the side of the machine’s reliability. It arises on the side of the organisation’s capacity to distinguish, a capacity that rested on indirect signals which abundance has just erased.

Almost no one has crossed that threshold. Hence the paradox observed in nearly every large research organisation. Very high usage rates, very high declared satisfaction, and no effect that anyone has been able to establish on what counts. Not on decision time, not on the length of the queue waiting in front of experimentation.

Programme success rate does not figure in that list and it cannot. Its latency is counted in years, no organisation collects the data, and I write in the fourth essay that it is what will settle the matter.

There is a simple way to check this diagnosis in your own house. Take your last R&D executive committee. Count the decisions taken. Compare with the one three years ago. If the number has not moved, and if the average time between a question emerging and its resolution has not moved either, then whatever your adoption rate, you have transformed nothing. You have given comfort back to people. That is estimable. But it is not what you budgeted for.

IV. Four Tiers, and Only One That Counts

I offer here a way of reading what is happening. Four tiers, which cannot be crossed out of order.

Tier 1 — Assistance. Everyone uses AI in their own way. The researcher to rephrase, the project lead to summarise, the manager to prepare a meeting. It is real, it is useful and indeed it is appreciated. It shows up nowhere in the company’s accounts and it never will, because the value is captured by individuals and dissipates at their level. It is an improvement in working conditions, and one ought to have the honesty to present it as such.

Tier 2 — Evaluation. The organisation equips itself with the capacity to say, for a given type of work, what distinguishes a good result from a bad one. Explicitly, and in a way that others can take up after it. This tier produces no direct value. It is a pure cost, a toll.

Tier 3 — Delegation. Once one knows how to score, one can hand over. A bounded perimeter, known criteria, traceability, a threshold beyond which a human takes back control. It is here, and not before, that economic value appears, because it is here that one stops adding work to people and begins taking it away.

Tier 4 — Redesign. The process itself is redesigned for an organisation where output is abundant and judgment scarce. One no longer does the same thing faster, one does something else. It is the tier of durable advantage, and there is no point talking about it before tier 2 has been crossed.

Here is the observation that interests me:

The great majority of organisations are at tier 1 and believe themselves at tier 3.

They believe it because they have figures. Adoption rate, active users, use cases launched, declared hours saved. All these indicators measure tier 1 precisely. None measures the crossing of tier 2. And since they rise nicely, they produce a sincere conviction of progress. They always rise nicely: declared satisfaction with a free and convenient tool is one of the easiest figures to obtain in the world.

The adoption rate is the indicator that could, unfortunately, cost you three years. Easy to collect, flattering, it rises on its own, and it is perfectly uncorrelated with the only question that determines whether your investment will produce anything at all.

The question is not how many of your people use AI. It is: on how many types of work can you say, without summoning your most respected expert, whether the result is good?

For most organisations, the honest answer is: not many.

Here is what the crossing looks like. A team decides to treat a single type of work, for example the review of the state of the art on a given question. A precisely bounded deliverable, and nothing else. It gathers three experts in the field and asks them to write down what they would require of a junior colleague on that deliverable. No one asks them to evaluate a tool. The conversation lasts four hours and it is painful, because they discover that they agree neither on the number of sources expected, nor on the treatment of contradictory results, nor on what to do with a literature that is largely Chinese and untranslated.

They end up producing a list. Eleven criteria, four of them disqualifying. They also write, which was not planned, a dozen examples of faults they hold to be fatal. They realised that it is easier to name what disqualifies than what qualifies.

Then they take fifteen past pieces of work, produced by humans, and score them with their own rubric. Three of those pieces, signed by respected people, fail the disqualifying criteria.

That is the moment that makes or breaks the project. Either the organisation treats those three failures as proof that the rubric is bad, and it returns to tier 1 having lost six weeks. Or it treats them as the first new information it has produced about itself in a long time.

Those eleven criteria will become eleven boxes, and it is better said before the house discovers it. That is what it does best, turning a criterion into a form in eighteen months, with no bad intent, because a form is filled in faster than a criterion is applied.

Only two things stand against it. The first has to do with what is written. A criterion that asks how the document treats the contradictory results published on the question, or what it has done with a literature that is largely untranslated, cannot be ticked by someone who has not read. A box of the kind compliance with the standard outline is ticked without reading anything at all. The difference lies entirely in the drafting of the criterion, and that is why this work falls to people in the field and never to a quality department.

The second is a test. A rubric that fails nothing for a year is dead. The one applied to the fifteen pieces of work failed three of them, signed by respected people, and that was the price of proving it measured something. The day that rate falls to zero, the rubric has become a box and there is nothing to save.

A written rule also lowers the cost of no. The one who refuses stops refusing in their own name and applies a criterion their peers wrote with them, which relieves them of having to give grounds alone and of the risk of being the one who has not understood the times.

What I can say is that the second case is rare, and that it is always the doing of a leadership that decided in advance that the results would be owned whatever they were. That is not improvised in the debrief meeting.

One last remark I wish to share here about this scale. It is crossed perimeter by perimeter, never organisation by organisation, and crossing one perimeter does not carry the next.

V. Why No One Pays the Toll

One might think tier 2 is a technical problem. Technically it is not very hard. It is tedious, methodical, and there is nothing in it that demands a breakthrough. What blocks is elsewhere. Building an evaluation of expert work means writing down in black and white what makes an expert right. And therefore, mechanically, when they are wrong. That is why it does not often get done.

For thirty years, scientific authority in these houses has rested on a functional opacity. The person who says no in four seconds has no need to justify, that is precisely the privilege their reputation bought them. Asking them to state their criteria amounts to asking them to make contestable what was not.

An explicit criterion applies to everyone. If you establish that a review must cite at least three contradictory primary sources, that criterion applies to the machine, and it applies just as much to the report your scientific director signed last month. You cannot build a rule that holds only for AI.

It is therefore, in the strict sense, a political act. It redistributes authority, from a category of people towards a set of criteria. It turns a personal power into an institution. It is exactly what organisations do when they move from interpersonal trust to rule, and historically, that has never happened without conflict.

There is an instructive precedent, and it is financial. For a long time, a company’s soundness was judged by the reputation of its head and the word of its banker. The move to standardised accounting, then to external audit, then to compulsory certification, took more than a century and met a certain resistance, because it took from directors the power to say what their company was worth. No one today would argue for returning to the banker’s word. Every step was fought by those whose latitude it reduced, with arguments that look a great deal like those heard in laboratories. That reality is too complex for rules, that the figure kills discernment, that the one who truly knows the file cannot explain what they know. Those arguments were not entirely wrong.

There is a second precedent, and I know it better because I come from it. Particle physics long watched discoveries announced at three standard deviations evaporate. The signal was real in the statistical sense, it did not survive verification, and the discipline took time to understand why. When you look in enough places, there is always a place where noise looks like a signal. It drew from this a single, deliberately brutal rule. Five standard deviations. One chance in three and a half million that noise alone would produce a deviation that large.

The formulation matters, and I correct it because the confusion is constant, including among those who invoke the threshold. What the threshold measures is the probability that noise alone produces what is observed. The inverse probability, that the observed deviation is noise, requires knowing what one believed before looking, and no one ever writes that down. It is exactly the inversion produced by a confidence score computed by the machine on its own answer.

The threshold is arbitrary and no one pretends otherwise. It renders a service that its arbitrariness does not diminish. It applies to the first-year doctoral student as to the Nobel laureate, and no reputation exempts anyone from it. This precedent is more encouraging than the accounting one, because the constraint was not imposed from outside. The experts gave it to themselves, on noting that their collective judgment produced too many errors to go on relying on it without a safeguard. Nothing forbids a scientific division the same move.

Hence the three evasions observed everywhere, which are three ways of going round the conflict.

The first consists in letting AI score itself. Attractive, quick, and it avoids going before the experts altogether. The system is asked for a confidence score on its own answer, and the decision is taken on that basis. This amounts to abolishing your internal control function on the grounds that the controlled party undertakes to be honest. No one would accept that for their accounts. Many accept it for their science, because the line between the one who produces and the one who judges has never been drawn in this field, there being no judge.

And the trap is more devious than it looks. An AI system is very good at producing a plausible confidence score. It does not contradict itself, it does not blush, it does not stammer. The score will look well calibrated, convincingly distributed, and it will have every statistical property of a measurement except that of being one. A measurement supposes a protocol fixed before observation and an uncertainty one knows how to estimate. You will obtain a magnificent dashboard.

The second consists in importing an external benchmark. One takes a public ranking, an available benchmark, and announces that the model is excellent because it scores well on it. It is comfortable. An external benchmark tells you whether a machine is good in general. It tells you nothing about what, in your house, makes a piece of work acceptable. Your criteria are particular. They carry the trace of your past failures, your industrial constraints, your historical arbitrations, sometimes an accident that cost dearly twelve years ago and that no one mentions any more but everyone takes into account. No one else has written them down, and no one else can write them for you.

The third, the most common, consists in referring the decision to the most respected expert. A pilot is built, the result is shown to them, they are asked what they think. They give an opinion. It is recorded as a validation. One has just reproduced, on a digital medium, exactly the reputation system one claimed to be moving beyond, with the added illusion of having advanced.

These three evasions have one thing in common. They make it possible never to hold the difficult conversation.

What makes it feasible, when it is, is an inversion of posture. The point is not to have a machine validated by your experts. It is to have them write the examination. Set the questions, define what a good answer contains, say what constitutes a fatal fault.

The difference is almost entirely psychological, and it changes a great deal. In the first case, someone is asked to give up authority. In the second, they are asked to fix it, to pass into a durable institution what they held personally and what will disappear with them. Many accept willingly when it is put that way. Some even experience it as a recognition, which objectively it is.

I add an observation, which I offer for what it is, a repeated impression rather than a record. The most reluctant experts are almost never the most senior. They are the mid-career ones, those whose legitimacy is acquired but not consolidated. The most senior, those who have nothing left to prove and are beginning to think about what they will leave behind, are often the best allies. The question of transmission already occupies them, and what is being proposed is exactly a mechanism of transmission.

One rule remains that admits no exception. The one who produces cannot be the one who scores. Not the machine on itself, not the team that built the tool on the quality of the tool, not the project sponsor on the success of the project. This separation looks obvious once stated, does it not?

VI. The One Who Checks Does Not Check

One belief remains to be dismantled and it has the appearance of common sense. We do not need to build an evaluation, since our experts review.

This is the default position almost everywhere. The human keeps control, the human validates, the human is in the loop. The formula has become so standard that it is written into usage charters without a second’s thought about what it covers. It covers this: an expert receives a document produced by a machine, skims it, and gives their agreement.

I believe that this review, in the great majority of cases, checks nothing at all. And I believe it is worse than nothing, because it produces a trace of control where there was no control.

Three reasons for that, and they are cumulative.

The first is that form disarms examination.

We have learned, over decades of professional reading, to use the formal quality of a text as an index of the quality of its preparation. A well-structured document, correctly referenced, free of language errors, with apt transitions and controlled vocabulary, signalled someone who had worked. This was not a prejudice. The inference was statistically founded, because producing that form was expensive and no one paid that cost for content they knew to be hollow.

The inference is dead. Form now costs nothing. The reflex is intact. It is too old, too solidly installed, too useful elsewhere to be switched off by conscious decision.

The result is that an expert reviewing a perfectly written text reads it with a level of vigilance calibrated on the old world. They look for defects where defects used to announce themselves: in hesitations, approximations of language, unsteady structures. They find none. Their vigilance drops. And it is then that they pass over the invented reference, which is written exactly like the other twenty-two.

It would be enough, so it is said, to warn people. They are warned. It is not enough. It is not information that is missing, it is a perceptual automatism, and a perceptual automatism is not corrected by a memo.

The second reason is that checking costs more than producing.

It is the most awkward inversion in this whole story. In the old world, producing a review cost two weeks and reviewing it cost two hours. The ratio was one to forty, which made the review economically obvious. Its usefulness was not even discussed.

Today, producing costs an afternoon of machine time. And checking properly, that is, going after each of the hundred references, opening it, confirming that it exists, that it says what it is made to say, and that it says so in the context in which it is invoked, costs far more time.

The ratio has inverted. Checking has become the most expensive operation in the chain.

No organisation consciously accepts that expense. So it does not incur it. It does something else, which it calls by the same name: it samples. The expert checks three references out of a hundred, chosen because they look unusual, finds them correct, and extends the verdict to the whole.

This reasoning would be legitimate if errors were randomly distributed. They are not. This is the difference between a statistical error and a systematic one. The first shrinks as the sample grows, the second never shrinks, however many references are checked. A generative system goes wrong precisely where the literature is thinnest. On lightly covered subjects, recent results, rare combinations. That is, on exactly the points of most value to you, since they are the ones where you are looking for something others do not have. Sampling has checked the trivial and let the decisive through.

That expense has three storeys and they do not carry the same price. The existence of a reference is checked by identifier resolution, at almost no cost and across the whole document. A few days of development suffice, and no expert need touch it. Knowing whether it says what it is made to say requires the full text, often behind a paywall, and a human reading. There remains the third storey, where one asks whether it says so in the context in which it is invoked, and that one delegates to nothing. It is the one that costs several days. An organisation that pays experts to do the first storey has the wrong expense.

The third reason is the most uncomfortable: the one who reviews does not want to find.

An expert asked to review a machine’s output finds themselves in an uncomfortable social position, and that position pushes in two opposite directions, neither of which leads to an honest check.

If they validate, they save time, they appear cooperative, they avoid looking like the brake on the project. And they shift the risk onto an impersonal object that will not be held against them. If they reject, they must give grounds, they delay, they expose themselves to someone else revalidating behind them, and they run the risk of being the one who has not understood the times.

In both cases, what determines their decision is not what the document contains. It is what their decision will cost them.

Add that reviewing is almost never paid, never measured, never valued in a career, and you obtain an activity everyone performs and no one has reason to perform well. We have built a quality control whose controller has no interest in exercising it seriously.

These three reasons reinforce one another. Form disarms vigilance, cost makes checking impracticable, the incentive pushes towards validation. The result is a mechanism that systematically produces agreements, and which, because it produces them systematically, carries no information.

A control that always approves is not a benevolent control. It is a control that does not exist.

I want to be precise about what that implies, because the conclusion is counter-intuitive.

It is not that one should review more seriously. That is the usual answer, and one does not push back against three structural forces with an appeal to professional conscience.

It is that human review becomes possible once the evaluation is built, and not before.

An expert asked whether this document is good is thrown back on their overall judgment, with no purchase, in the conditions described above. They will validate. An expert asked whether this document satisfies these four disqualifying criteria, of which here are the definitions, does an entirely different job. They have a grip, they know where to look, their decision is driven by something other than their social position, and any rejection is no longer a personal opinion but the application of a rule they themselves helped write.

It is the same person, the same document, and two operations with nothing in common.

One sees then that asking whether a human is needed in the loop teaches no one anything. The answer is obviously yes. What must be asked is: with what in their hands.

A human in the loop without an instrument is no guarantee. It is an alibi, and one that organisations grant themselves all the more readily because it is free.

VII. Evaluating Kills Exploration

The opposite camp has a serious argument: evaluating kills exploration.

It is most often put by the most creative researchers. An organisation which scores everything ends up producing only what is scorable. Explicit criteria reward conformity and penalise deviation. The discoveries which count almost always have the appearance, at the moment they occur, of a poorly founded result. An ill-supported hypothesis, a lead that contradicts the literature, an intuition one cannot yet justify. If you had submitted those pieces of work to a rubric, the rubric would have rejected them.

That is true. Historically, it is massively true. The history of science is full of work that would have failed any criterion of its time, and the history of research evaluation, bibliometrics, impact factors, assorted indices, is a catalogue of documented perverse effects. When you measure publications, you get publications. When you measure citations, you get citation networks. The principle is known and it is sturdy. Any measure that becomes a target stops being a good measure. It is Goodhart’s law, stated originally for monetary policy and verified since wherever steering by indicator has been attempted.

I add the aggravation that comes with AI: a system trained on what exists is structurally conservative. It produces consensus fluently and deviation reluctantly. If you build an evaluation that rewards conformity to the literature, you will get a machine excellent at repeating to you what you already know, and you will call that a success, because your scores will be superb.

So the objection holds. It simply does not hold where people think.

It does not say that one should not evaluate. It says that there is no single evaluation, and that everything depends on how many instruments one has. There are two natures of intellectual work, and they do not obey the same control.

There is work that must be true. A review of the state of the art, a compilation of results, an evidence file, a regulatory note. Here the requirement is binary and merciless. An invented reference disqualifies the document, it is not weighed against other criteria. An evidence file of which not one line can be guaranteed is not worth less than a good file, it is worth less than nothing, because it consumes verification time it claimed to save. On this type of work, traceability is the condition of entry.

And there is work that must be fertile. A hypothesis, a lead, a reformulation of a problem, an analogy from elsewhere. Here, applying the criterion of truth is a category error. A hypothesis is neither true nor false at the moment it is put forward. It is plausible, original, and above all testable. Three demanding criteria, none of which resembles conformity. The question is not whether the literature confirms it. It is whether one could build the experiment that settles it.

Testability penalises vagueness, but it rewards deviation: a hypothesis that conforms to the literature is generally of little interest to test, since the answer is already known. A rubric that scores testability therefore pushes towards boldness, not towards conformity: provided it is not contaminated by criteria of truth that have no business being there.

There is a third nature of work, which I discuss less because it is less discussed, and which is the most frequent in an industrial organisation. Work that must be conclusive. A recommendation for arbitration, a decision-support note, a comparison of options. This work bears on an uncertain future, so it does not have to be true in the strict sense, and it does not have to be fertile either. It has to be conclusive. Its criteria are different again. The exhaustiveness of the options considered, the making explicit of assumptions, the resilience of the recommendation to variations in those assumptions. A note that recommends A and would still recommend A if three parameters were changed is a good note. A note that swings to B at the slightest adjustment is a bad note, however impeccable its reasoning.

Three natures, three rubrics. An organisation that applies the rubric of truth to work of fertility sterilises its research. An organisation that applies the rubric of fertility to work of truth poisons its corpus. An organisation that applies either to conclusive work produces brilliant notes no one can do anything with.

I therefore concede the objection entirely, and turn it round. The risk is not evaluating, it is evaluating with a single instrument. That is exactly what the reputation system does. A single instrument, trust in a person, applied indifferently to everything.

A second objection remains, shorter, and it is often the real one: it is a great deal of effort.

Yes. Building a serious evaluation on one type of work is a few person-weeks of experts who have other things to do. It is a real cost, and it is visible, whereas the cost of doing nothing is invisible.

But look at what you are buying. Not the ability to score a machine, the right to delegate anything to it at all. And, incidentally, the setting down in writing of a knowledge of judgment that until then existed only in the heads of a few people and left with them. Many organisations discover, in building their evaluation, that what they have chiefly done is document their own expertise for the first time in their history.

The side effect is perhaps the main benefit, and it would have justified the effort even without AI.

VIII. Back in the Room

Let us return to the meeting room.

What happened there that Tuesday was not an adoption incident, nor resistance to change. That woman did not resist. She did what she had been asked to do for thirty years, arbitrate with her judgment, in an organisation that never built an alternative.

The problem was not her. The problem was that the whole room depended on her and did not know it.

That is what AI is revealing, and that is why this subject will not stay technical for long. An organisation that cannot evaluate its expert work is an organisation whose quality rests on a small number of irreplaceable, undocumented people. This was already true before. One could avoid seeing it.

The whole of this text can in fact be stated without ever mentioning artificial intelligence, and that is a good test of its soundness. An organisation that cannot say why a piece of work is good can neither train, nor transmit, nor settle a disagreement, nor defend itself before a regulator, nor survive the departure of its figures. All of that was true in 2015. AI adds nothing to the finding. It only removes the possibility of ignoring it, because it places before you, every day, dozens of outputs about which no one can say whether they are worth anything.

And the most important point, the one I believe leaders underestimate entirely: they think they are arbitrating a technology budget when they are arbitrating where authority resides in their house, in people or in criteria. It is a matter of internal constitution, not of tooling. It arises perhaps once a century in an organisation. It is arising now, through the side door, in steering committees where the talk is of licences and usage rates.

Those that do not cross tier 2 will not collapse. It is more insidious than that. They will stay at tier 1 for years, with excellent adoption figures, real internal satisfaction, and no improvement in their capacity to decide. They will have spent, they will have communicated, and they will be able to delegate nothing, because one does not delegate what one cannot control. They will notice the day a competitor has decided three times faster than them for two years, and on that day the accumulated gap will no longer be closable by investment.

Those that cross it will first pay an uncomfortable price. A few slow months, unpleasant conversations, experts who must be persuaded to write down what they knew without saying, and at least one moment when the rubric says something no one wanted to hear. Then they will hold something no technology buys, the ability to say this is good without summoning anyone. From there, on that perimeter, they will be able to delegate and to redesign. On the neighbouring perimeter, they will start again from zero. That is the price of this business.

To measure where you stand, no tool is necessary.

Take the last important decision made in your R&D, the stopping of a programme, the choice between two technical routes, the allocation of a significant budget. Ask yourself whether, today, someone new could reconstruct why it was taken. Not the minutes of the meeting, the reasons. The options set aside and the grounds for setting them aside. The underlying assumptions and what supported them.

In most cases, the answer is no. What remains is a decision, a date, signatures, and a knowledge that now exists only in a few memories. Which, in ten years, will be elsewhere.

The problem goes beyond archiving. Reasoning was never treated as a deliverable. The conclusion was delivered. The path stayed with the one who had walked it, and no one found that abnormal, because the one who had walked it would be there to walk it again if needed.

That assumption, they will be there, is the one that has just fallen, and it has fallen for two simultaneous reasons. Demography first. The generation that holds this knowledge in these houses is largely at the end of its course. Abundance next. Even if it remained, it could no longer arbitrate alone a volume of output multiplied by a hundred.

It is the conjunction that makes the moment particular. Each of the two causes, taken alone, could have been absorbed. Together, they shake the organisation.


I end on a remark that is not reassuring but seems to me right.

An organisation rarely crosses tier 2 because it has understood an argument. It crosses because a precise event forces it to: a file rejected by a regulator for an unverifiable reference, a programme stopped too late that cost two years, the simultaneous departure of three experts in one field. The trigger is always a loss, never a lucidity.

There is nothing surprising in that. Tier 2 is a certain cost against a probable benefit, which is the kind of arbitration organisations make worst. An outside force is needed to make it acceptable.

What I can say is therefore modest: that force will come. It takes the shape of a regulator, a competitor or a retirement, and its timetable is not yours to set. The only variable that is, is whether you will have started before it arrives.

The real AI project in expert organisations is therefore not an AI project.

It is that of the share of judgment: how much of it is left, where it sits, and to whom it belongs.

Be notified when the next essay appears.

Your address is used for this message and nothing else. It is neither sold nor rented to third parties, and can be withdrawn in one click. Privacy policy.

Thank you. You will receive a message when the next essay appears.