Essay II
The Expert Afterwards
On what we were paying for, without knowing it, in those who know
You had reached section
I. The Wrong Fear
The question always comes, and it comes quickly. As soon as one explains that an organisation will have to learn to evaluate the work of its experts, someone raises a hand. And them, what becomes of them?
I will answer frankly, because the soft answer usually given, AI does not replace, it augments, is true but at the same time perfectly useless next to the point being discussed here.
Your experts will not be replaced, let us be reassured on that. But half of what you were buying from them has just lost most of its value, and it is precisely on that half that you recruited, evaluated, promoted and paid them for several decades.
The fear is therefore misplaced. AI will not take your experts’ work. The founded fear lies elsewhere: I never properly worked out what I was paying for in them, and I am going to have to learn it now, urgently.
Let me say at once what this text is not. It is not a prediction about the number of jobs. Such predictions are unverifiable, the history of the last thirty years has shown that they are rarely accurate and often rather sensational. It is, rather, an observation about the internal structure of the expert’s trade, about what it actually contains, about what of that content has just been absorbed by a machine, and about the fact that our talent management systems measure exactly the absorbed part.
This observation is not reassuring. It is more useful than a prediction, because it bears on decisions you can take as early as this year.
II. Opening the Block
Expert work is spoken of as a block. It is one for whoever observes it from outside, and it is one of the rare trades where that opacity has been accepted without ever being lifted. She is an expert, she does expert work.
This opacity is not accidental. It is constitutive of the learned professions, which were all built on the same movement. Delimit a domain, control access, refuse those outside the right to judge what happens within. That is what distinguishes a profession from a trade. The doctor, the lawyer, the researcher have in common that they obtained that only their peers could evaluate their work. In exchange, they collectively guarantee the quality of what their members produce.
The market long accepted this arrangement because it had no choice. There was no other way of knowing whether the work was good.
Let us open the block, since we now have the means. It contains four distinct gestures.
Collecting. Gathering what exists. Reading the literature, retrieving internal results, querying databases, identifying who has already worked on the question, reconstructing the history of a file. It is patient work, time-consuming, and it occupies a share of the time no one dares measure because the figure would be awkward.
Producing. Turning that material into a readable object. A synthesis, a protocol, a file, an analysis note, a presentation. This is where visible competence plays out. Clarity, structure, command of vocabulary, the ability to make the complex intelligible. It is also, very largely, what an expert is judged on by their peers.
Arbitrating. Choosing. Deciding that this lead is worth more than that one, that this result is sound and that one doubtful, that one should stop here rather than continue, that a weak signal deserves attention. This gesture produces nothing visible. It leaves no trace. It is practically never documented, and yet it determines whether the next six months will be useful or lost.
Committing. Owning it. Putting one’s name to a recommendation, defending it before a committee, bearing the consequence if it is wrong. It is a gesture of responsibility rather than of competence, and it is the only one of the four that no one today disputes will remain human.
Four gestures, one trade, and a distribution of time that is roughly the inverse of the distribution of value. Depending on the field, collecting and producing absorb between sixty and eighty per cent of a senior expert’s time. Arbitrating and committing occupy the rest, and it is the rest that justifies their salary. This imbalance very often passes for normal.
No one would find it normal for a surgeon to spend seventy-five per cent of their time preparing the theatre. The operating theatre was organised precisely to avoid it. Around the surgeon, a whole chain of people does everything that is not the operation, because their time is the scarce resource and it would be absurd to consume it otherwise.
In research, that was never done. The expert collects, formats and reviews personally. This was not an organisational choice. Those tasks were inseparable from judgment, and one could not delegate the reading of the literature to someone who would not have known what to look for in it. The low gesture was the means to the high one. It is that inseparability that is now coming apart.
The human collector had to know what to look for because he had to give things up. He could not read a hundred articles, he kept twelve, and the sorting happened during the collecting, inside a single head, without ever being spelled out. The machine does not give things up. It hands you the hundred (two of which do not exist, incidentally), and the sorting that was buried inside the low gesture surfaces as a separate operation. It is the same judgment as before. It has changed place, and it has become visible now, which is convenient for whoever wants to evaluate it.
The time freed up therefore arbitrates on material that is more abundant and less filtered than before. That is harder, and it takes up more than the hours saved.
III. The Line of Fracture
Generative AI does not cut the block just anywhere. It cuts it between the second and third gesture, and fairly cleanly.
Collecting: absorbed almost entirely. It is observed everywhere, and it is the first use experts adopt of their own accord without being asked. One important reservation remains. The machine does not go looking for what is not written, and a considerable share of an organisation’s knowledge has never been written. On everything published, indexed, accessible, human collection no longer has an advantage, and there is consensus on this point.
Producing: absorbed for formatting, largely absorbed for structuring, not absorbed for correctness. The machine writes better than most researchers. It is a sentence that may displease but that is largely true. The average writing quality of a text produced by a good system exceeds that of an internal note. What the machine does not know is when it is wrong.
Arbitrating: not absorbed, and probably not absorbable in the short term. The machine may perhaps be capable of it in the abstract, but arbitrating supposes knowing what is not written. Your past failures. Your industrial constraints. The history of that programme stopped a few years ago and the reasons no one wants to return to it. The relative reliability of such a supplier, such a partner laboratory, such a type of measurement. None of that exists in written form, and none of it will be written spontaneously.
Committing: out of reach, and for a long time. One does not delegate responsibility to something that can lose nothing. It is a legal obviousness, and it is also an anthropological one. We grant credit only to those who take a risk in speaking.
There is the line, and it is fairly clean. It separates work whose value comes from effort from work whose value comes from implicit knowledge. The two absorbed gestures are exactly the ones your management systems measure.
Look at what you recruit an expert on. A degree, publications, a doctorate, in other words proofs of a capacity to collect and to produce. Look at what you promote them on. The volume and quality of their output, seniority, possibly external recognition. Collecting and producing again. Look at what you pay them on. A seniority scale that has never contained a single line about the quality of their arbitrations.
Let me put the question differently, because this is where it becomes concrete. On what document, in your organisation, is it written that this person takes better decisions than that one? You have built an entire organisation around the half of the trade that has just been automated.
Annual reviews do contain appreciations of discernment, vision, decision-making capacity. That is correct. They are always qualitative headings, phrased in adjectives, without an example, without reference to a specific decision and to what became of it. They are not false. They are not backed by any observation. No one, in those reviews, says last year you recommended stopping programme B, you were right, here is what it saved. No one followed what became of the recommendation.
IV. The Invisible Downgrading
What happens next is predictable, and widespread enough that one should stop seeing it as a particular case. A senior expert uses AI. They gain time, really, there is no illusion about that. They do something with it. They produce more. More syntheses, more notes, more files.
They therefore produce more of the thing that no longer has value. Volume rises. Formal quality rises. The expert’s satisfaction rises, because they do faster what they liked doing. Nothing alerts them.
Meanwhile, the only thing that should have increased has not moved an inch: the quantity and quality of the arbitrations delivered. They arbitrate the same number of times as before, often less well because they arbitrate on more abundant material and have no more time in which to do it.
This last point is counter-intuitive, and important. One spontaneously imagines that more information improves the decision. That is true up to a certain volume. Past a threshold, additional information no longer reduces uncertainty. It increases the sorting load, it multiplies the considerations to be weighed, and above all it produces a side effect: an increase in confidence without an increase in correctness. Someone who has read forty pages on a question feels better informed than someone who has read twelve. They do not decide better. They simply decide with more assurance.
And one day, someone notices that an equipped junior produces a document roughly equivalent to theirs. That is false, the senior’s document is better, but the gap has moved from considerable to arguable, and arguable is enough. That day, the legitimacy that rested on output begins to crack, and the expert has nothing else to set against it, because the part of their trade that could have defended them was never named.
That is what I call, here, the invisible downgrading. It is observed retrospectively, at a reorganisation, when someone wonders aloud why three seniors are paid for this activity.
The most troubling thing is that the expert takes part in their own devaluation, and with enthusiasm. They adopt the tool, they promote it, they sometimes become the internal reference. The tool makes the work more pleasant for them, and being ahead on the use of a technology has always been a positive marker. They are demonstrating, zealously, that the visible part of their trade can be done by something other than them.
I am not saying they should abstain. I am saying that no one around them helps them see what is at play, because the organisation itself does not see it.
A lucid organisation does the opposite: it stops measuring what its experts produce and starts measuring what they set aside.
Simple to state, and yet: no one does it, and not through negligence, because setting aside leaves no trace. An expert who has avoided six months of useless work in ten minutes of conversation appears on no dashboard. They will appear the day they have gone and no one avoids them any more.
There is an asymmetry here that structures the whole subject. Errors committed are visible, errors avoided are not. A decision that goes badly produces an observable consequence, a cost, often an enquiry. A decision that averts a disaster produces nothing at all, since the absence of disaster is not an event. Every organisation therefore overestimates the cost of bad decisions and underestimates the value of good ones.
That asymmetry was bearable when judgment was abundant. It becomes dangerous when it becomes the scarce resource.
V. The School We Are Closing
I come to the objection, and I must say at once that I find it strong.
Judgment is not taught. It is acquired by producing.
No one learns to spot a doubtful result in a course. One learns it having produced two hundred results oneself, of which thirty were doubtful and four blew up in a meeting. One learns to sense that a lead will go nowhere because one spent eighteen months on a lead that went nowhere. One learns to distrust a measurement that is too good because one once built a whole demonstration on a measurement that was too good and was false.
Judgment is sedimented failure.
The ten years of thankless production that produced that sedimentation are precisely what AI absorbs. Today’s junior will not do the tedious collection, will not write the forty mediocre syntheses before the first good one, will not be wrong often enough to develop the instinct.
We are therefore closing the school of judgment at the exact moment judgment becomes the scarce resource.
I add that the phenomenon is not confined to research.
Aviation is the documented case. An airline pilot hand-flies almost all of their flights and goes back into a simulator under check every six to twelve months. What degrades is manual handling at cruise, when flying is left to the autopilot. The degradation is serious enough that the American authority had to recommend, twice, in 2013 and again in 2017, that pilots hand-fly more. The pilot whose hand has grown unaccustomed does not know it.
Medical imaging showed the same degradation back in the diagnostic-aid era of the two-thousands, and then recent systems brought detection back up. The deskilling signal is observed today mostly in digestive endoscopy. As for the finance of 2008, I withdraw it from my list. It is explained by incentives and by the correlation of distribution tails, not by a deskilling of risk managers.
The average performance of an automated trade can rise while recovery competence falls. It does not show in the averages. It shows in the tail of the distribution, at the moment of crisis, and never before.
The objection is therefore rather good, and it contains a premise worth examining. The premise is: it is by producing that one learns to judge.
Let us look at what, in production, actually teaches. It is not the gesture of producing — the thousandth synthesis teaches nothing the hundredth had not already taught. It is the confrontation with error, and more precisely the feedback on a decision one has taken.
That feedback was frankly slow, and it was slow for purely material reasons. A researcher saw their errors sanctioned twice a year at best, sometimes every three years if the experimental cycle was long. That is why the formation of judgment took fifteen years. It was not that much had to be produced, it was that one had to wait a long time.
That distinction changes a great deal. If the limiting factor was the volume of production, then removing production destroys the training, and the objection is decisive. If, on the other hand, the limiting factor is the delay of the feedback, the question becomes quite another: how does one manufacture fast feedback loops on decisions whose sanction is slow?
That is a different problem, and it is a problem one can work on. The phenomena a researcher studies are irregular; the ways of being wrong about them are much less so. Mistaking a correlation for a mechanism, stopping at the first sufficient explanation. These faults are recognisable ten years later, in another field and on another literature.
Confronting a junior with three hundred past arbitrations whose outcome is already known, and asking them to decide before showing them the reasoning that was held at the time, and then what happened. Systematically bringing back, two years later, what became of each choice, which no one does though the information exists. Documenting rejections and not only acceptances, because that is exactly where the knowledge sits: an organisation keeps all its positive decisions and, too rarely, its negative ones, whereas it is the second that contain the judgment.
I do not claim that this reproduces fifteen years in the job. I say that the formation of judgment is not designed. It is an accidental by-product of slowness. An accidental by-product can be replaced by something deliberate, and there is even a chance the deliberate does better, because no one has ever tried, and because an arrangement that concentrates three hundred decisions into six months offers a density of learning no natural career has ever offered.
There is a precedent, and it is interesting because it concerns exactly a competence of judgment under uncertainty. Airline pilots no longer learn emergency situations by living through them. They learn them in a simulator, where they can be taken through more failures in a day than a whole career would produce. No one can properly compare the two kinds of learning. It would take a control group put through real failures, and that group will never exist. What is measured is more modest: in the simulator, rare situations are deliberately provoked, performance on them is recorded, and it starts over until it holds. A whole career does not produce enough failures to train anyone to handle them.
The price of this arrangement is half the lesson. Currency of competence in aviation is mandatory and regulated, backed by a type qualification and a whole apparatus of checks, and it represents a permanent expense the industry has borne for decades. Confronting a junior with three hundred past arbitrations costs infinitely less, and is worth infinitely less. I am proposing a starting point, not the aeronautical equivalent.
There remains the part of the objection I do not know how to treat, and I prefer to state it rather than go round it.
Part of expert judgment is of the order of what Michael Polanyi called tacit knowledge, of which he said that we always know more than we can tell. The discomfort before a figure that does not fit, without being able to say why. The feeling that a line of reasoning is unsteady before having identified where. That part is real. It is probably irreducible to any formalisation, since by definition it escapes the one who possesses it. And I know of no arrangement that transmits it other than by long companionship, that is, by prolonged presence alongside someone who has it.
If that part is the greater one, the objection wins and I am wrong. I believe it is the smaller one. The greater part of what is called expert intuition seems to me to be unwritten contextual knowledge, which could be written if one took the trouble. I cannot prove it, and whoever asserts the contrary cannot prove it either.
That is the honest zone of uncertainty of this text. It deserves to be treated as such rather than covered over with conviction.
VI. The Most Costly Reflex
The reflex, at this stage, is to summon training. A plan, modules, an acculturation, a platform. It is the most costly reflex there is, and it treats as a competence problem what is a structural one.
The subject is not training your experts in AI. It is changing what you buy from them, and therefore what you measure, what you promote, what you pay for. As long as those three systems reward output, no training will shift anything. People do what they are evaluated on, and they are right to.
Three decisions can be taken, in this order.
1. Make arbitrations visible.
Today, an expert decision leaves a deliverable and no trace of the decision itself, neither of the options set aside, nor of the reasons, nor of the assumptions that founded it. This must be inverted, and the rejection documented as much as the choice.
Nothing to do with a governance project or with bureaucracy. It is the condition for anything else to become possible. One can neither measure, nor transmit, nor evaluate what is not written. And the cost is low if one resists the temptation of the form. Three lines at the moment of the decision, what I set aside, why, what would make me change my mind, are worth more than an eight-page template no one will fill in.
The third line is the most important and it is the one always forgotten. What would make me change my mind is the only formulation that makes a decision verifiable after the fact. Without it, one cannot know whether the person was right or lucky.
2. Separate the two careers.
The dual ladder already exists, and I am not going to present it as a new idea. American industrial laboratories invented it at the turn of the 1950s, and it is in place in most of the houses I am speaking of, under the names of expert, fellow, research director as opposed to research engineer. The technical track becomes the shelf where those one does not want to promote are put away.
The existing expert ladder measures technical depth and is earned by producing, which makes it a second production ladder. There is really only one ladder that counts, then, and it mixes output and judgment. One progresses by producing a great deal, then one day finds oneself arbitrating without anything ever having checked that one was capable of it. It is a design error that was without consequence as long as the two gestures came together, and that ceases to be.
What I propose is orthogonal to the expert ladder. The separation I suggest bears on the promotion criterion, and on a single one. Has this person recommended things that turned out to be right, and has it been checked. A judgment ladder that borrows the criteria of the expert ladder will reproduce the shelf, and faster, because this time no one will even be able to invoke the publications.
There remains the main cause of failure: a track that carries no power over resources is a dead end. If the judgment ladder does not carry the right to stop a programme, it will be worth nothing, and it would be better not to open it at all.
Two trades, two trajectories, two systems of evaluation. This is about recognising that an excellent producer is not necessarily a good judge, that the reverse is equally true, and that an organisation which does not distinguish the two promotes the first into the posts of the second.
3. And the most important point, decide what you do with the absorbed half.
You have a choice between two paths, and refusing to choose amounts to choosing the first by default.
The first consists in recovering the time gained as a reduction. It is legitimate, rational, and produces a measurable result the following quarter, and there are situations where it is a responsible choice. It has a cost no one will put in the business case. It removes the stock of people in which tomorrow’s judges were to form themselves, and it sends those who remain a perfectly clear signal about what they are.
The second consists in reinvesting the whole of the time gained in arbitration, more decisions, taken earlier, better instructed, with feedback loops built on purpose. It produces no visible result for two years. It is very hard to defend before a committee that expects figures.
I am not going to claim that the second is obviously the right one. I distrust texts that always recommend the long-term investment from a position where they do not have to bear its cost.
What I do say is that the first is almost always chosen without having been decided. It settles in by default, through an accumulation of budgetary micro-arbitrations none of which ever mentions the word expertise. A departure not replaced here, a recruitment postponed there, a hiring freeze that was only to last a year. No one signs that decision, and everyone endures it. And five years later, one finds there is no longer anyone capable of saying no in four seconds.
VII. The Junior Without a Ladder
There is a consequence of all this that is discussed more and more within the community, and that seems to me essential because it is irreversible.
The structure of an expert organisation is pyramidal, and that pyramid has a function the org chart does not state. It produces competence. Juniors do the base work. In doing it, they learn. In learning, they rise. The seniors who leave are replaced by those who have spent fifteen years climbing. The system reproduces itself.
That system had a property no one ever needed to name. The base work was simultaneously an economic output and a training arrangement. The organisation paid juniors to produce, and it obtained as a bonus, without having budgeted for it, the next generation of seniors.
It is one of the rare mechanisms where short-term interest and long-term interest coincided perfectly. No one had to arbitrate between the two. No one therefore had to worry about preserving the second.
AI is decoupling them. The base work can now be produced without a junior. The economic output is preserved, improved even. The training arrangement disappears, and it disappears silently, because it never figured on the balance sheet.
Look at what happens concretely in a practice, a laboratory, a research department. Fewer juniors are recruited, because fewer are needed. Those recruited do something else. They supervise, they check, they orchestrate. This is called moving up the value chain, and it is presented as good news. They are spared the thankless work, they are put earlier onto higher-value tasks.
Except that they are asked to supervise work they have never done.
A junior checking a synthesis produced by a machine has none of the footholds that made checking possible. They never looked for those sources, so they do not know which ones should have been there. They never wrote that kind of text, so they do not know where one cheats. They were never caught out, so they have not developed the distrust that serves for it.
There are measurements on this exact point, and I must mention them here because at first reading they do not go my way. Recent work on generative assistance in business, in customer service, professional writing, consulting and software development, finds the same thing. Gains are largest among the least experienced, and the performance distribution tightens from the bottom. Their authors see this as the diffusion, to novices, of the tacit practices of the best.
That measures task performance under assistance. What I am talking about is what remains once the assistance is removed, and the two claims sit together without conflict. A junior who produces at a senior’s level with the tool has had their output raised, and nothing says they have learned anything. The second finding of that same literature is less cited and it does go my way. A degradation is observed when the assistant is withdrawn, and the best-established case concerns a gesture of examination, the one I cited above. None of these studies measures what occupies me here, which is the capacity to know which sources should have been in a document and are not.
They have been given the control function while being deprived of what made it exercisable. And they know it, which is what makes the situation particularly uncomfortable for them. Many young professionals today describe exactly that sensation. Producing, very fast, things they would not know how to produce themselves, and having to answer for them.
There is a learning problem here, and there is also a problem of implicit contract.
The implicit contract of an expert trade has always been this. You accept ten years of hard work, poorly paid relative to your qualification, and in exchange you obtain at the end an expertise that makes you rare and that no one can take from you. It is a considerable personal investment, and it was justified by the value of what one acquired.
What becomes of that contract if the hard work disappears but the expertise is no longer acquired? It becomes this. You are employable immediately, you are interchangeable immediately, and you will remain so. This contract is not bad for everyone, it suits those who did not want the long commitment. It is catastrophic for the organisation, which no longer manufactures anything.
I hear the objection, and it is reasonable. This has happened before, trades recompose themselves, the same was said of mental arithmetic and logarithm tables.
It is reasonable but it misses a point. What has been automated so far, in expert trades, were operations whose mastery was not constitutive of judgment. Knowing how to extract a square root by hand teaches nothing about how to choose a model. The slide rule disappeared without damage because it was only a tool.
What is disappearing today is not a tool. It is the repeated experience of confronting the material and being wrong about it.
A test, an unpleasant one, to know whether your organisation has this problem. Ask yourself where, in ten or fifteen years, the person who will know how to say no in four seconds will come from. Name them. If you cannot point to anyone under thirty-five today of whom you think they will be that person, then you have already stopped producing what you live on, and you stopped without the decision appearing anywhere.
VIII. The Price of No
One question remains that I have only touched on and that will end up being asked everywhere. If judgment becomes the scarce resource, what is it worth?
The economic answer is simple in theory. A rare and indispensable resource is paid for dearly. Its remuneration should therefore rise sharply, and the gap between those who have it and those who do not should widen.
That is what happens in every market where scarcity is observable. It is not what happens here, and the reason lies in the shape remuneration takes in these houses. A grid-based system can only pay for what it can write in an opposable box, and the quality of arbitrations does not appear there.
The economy knows how to pay for qualities it does not measure. The tournament, deferred promotion, partnership in a firm, the discretionary bonus that holds through repetition. None of these arrangements requires a measurement, all of them require only that one be able to rank two people after the fact. The instruments exist but run empty for want of the ranking they presuppose.
Look at how the market for expert labour sets prices today. It uses proxy indicators. The degree, seniority, the names of previous employers, the number of publications, the size of the teams led. All of these measure accumulation, not discernment. They worked because they were correlated with discernment. Someone who had produced a great deal had generally judged a great deal.
That correlation is breaking. One can now produce a great deal without ever judging. The indicators will therefore go on working in appearance, while measuring less and less of what they claimed to measure, and no one will notice quickly, because there is no competing indicator to compare them with.
What is going to happen, I believe, is a divergence between two markets.
The first will stay visible, formalised, structured by scales and competence frameworks. It will go on paying for output, with constant downward pressure, since output is becoming abundant. It is the market where one recruits on a CV.
The second will be informal and settled by private agreement. It will concern a small number of people known, in a given sector, to be right more often than the others. That knowledge already circulates, through recommendation, reputation, the call to a former colleague. It is written nowhere, it is remarkably reliable for the few dozen people who have access to it, and it carries no weight before a remuneration committee, because a signal that cannot be laid on a table does not enter a negotiation. On that market, prices will rise, and they will rise with no relation to the scales.
Real scarcity leaves a trace in prices. Four signs are observable without privileged access. The dispersion of salaries within a single trade and a single level. The frequency of counter-offers at the moment of departures. The day rate of independents whose reputation for judgment is established. The share of the external-expertise budget that goes to advisory assignments rather than production assignments.
I have no series on any of the four, and I know of no one who keeps one. What I am putting forward is therefore a prediction.
Organisations that see only the first market will make real savings on their headcount and pay considerable sums, in consulting and external expertise, to buy back at a high price the judgment they will have stopped producing. Many already do, without making the link between the two movements.
The question puts itself more directly: what is someone who says no worth, in your house?
Not someone who blocks, systematic blocking is a fault rather than a competence. Someone who, two or three times a year, stops something everyone wanted to do, and is right to stop it.
Half of their economic value is calculated without difficulty. A programme stopped at eighteen months leaves the budget a known line, the teams no longer mobilised and the trials no longer commissioned. Do the sum once, on a programme stopped at your own house last year, and set the amount against a senior’s fully loaded salary. The ratio is rarely counted in single units.
The other half cannot be calculated. What the programme that was killed would have delivered, no one will ever know, and the estimate always tilts in favour of the person who said no, since the stopped project never produces the proof that it would have succeeded. In a portfolio where most programmes fail, systematic blocking would in fact obtain considerable apparent correctness at zero competence.
There is only one way to make the subtraction honest: write down in advance, before knowing the outcome, what would have been launched without that opinion, and what would count as being proved wrong.
No organisation does it, and yet the calculable half costs only an afternoon. The difficulty is not technical. This calculation would make explicit a hierarchy of value the organisation prefers to leave implicit, because making it explicit would force salary and status consequences that would jar with the whole rest of the edifice.
Here we find exactly the mechanism of the previous text. To measure judgment is to make it contestable, comparable, and therefore negotiable. That is why it is not measured, and that is why it is not properly paid.
The irony is that this refusal protected experts as long as their legitimacy rested on something else. It exposes them far more now. A profession that refused for a century to make its judgment measurable finds itself with nothing to set against a demand to show what that judgment is worth, at the precise moment it would need to. The knowledge exists and it circulates. It stops at the door of the bodies that set the prices.
That is the backlash of opacity. It long served as protection. It has become a vulnerability today.
IX. The Manager Who No Longer Knows Who Is Good
There is a practical and immediate consequence: managers of expert teams are losing their ability to tell their people apart.
This deserves precise description, because it is a phenomenon occurring now, in nearly every organisation concerned.
How did a team leader know, until now, who was good? Not through annual reviews, which are a ritual of after-the-fact justification. They knew it through a continuous and roughly free signal, the quality of what landed on their desk. A clear, well-constructed note that anticipated their questions signalled someone who thought well. A confused note signalled the opposite. The signal was noisy, sometimes unfair to those who wrote badly and thought correctly, and it carried a great deal of information.
Today, everything that lands on a manager’s desk has become well written. Structured. Complete. Anticipating objections. The distribution has narrowed to the point where reading a deliverable teaches almost nothing about the person who produced it.
Evaluation falls back on what remains observable. Presence, availability, ease in meetings, the ability to get noticed. Social markers that never had much to do with the quality of the work, and that systematically favour the same profiles. Twenty years were spent trying to objectify professional evaluation precisely to reduce those biases. The objectification has just been removed while the biases were kept.
Meanwhile, real gaps widen while visible gaps narrow. A colleague who uses these tools with discernment, who knows when not to believe them, who checks what must be checked, who employs them to explore rather than to conclude, produces work whose value is far above that of the colleague who employs them mechanically. But the two deliverables look alike. The difference shows only downstream, when one of the two leads to a decision that holds and the other to a decision that breaks, that is, eighteen months later, when no one makes the connection any more.
The most perverse part lies elsewhere. The manager can no longer rely on their own judgment as a reader, because they are subject to exactly the mechanisms described above. The well-formed text disarms their vigilance too. They validate faster, they probe less, they trust more. They become, without noticing, a worse manager of quality at the moment quality becomes harder to establish.
Organisations are doing nothing about it, for now, because the problem is not named. When they take it up, I fear they will do the most predictable and most counter-productive thing: forbid or track.
Forbid the use of the tools on certain deliverables, to restore the signal artificially. That is untenable and will be circumvented in three weeks. Or else track usage, measure who uses what, how often, on which documents, to try to reconstitute information about people. That is equally useless, because intensity of use says nothing about quality of use. The most discerning colleague may be the one who uses them most.
The only way out I can see is the one this text has been describing from the start, and it is consistent with all the rest: if the deliverable no longer says anything about the person, one must look at the decisions.
No longer is this note well made, but has this person, over the last twelve months, recommended things that turned out to be right. No longer the quality of the document, but the quality of what it produced in the world.
That is a far harder evaluation to operate. It requires following recommendations over time, which no organisation does on its research decisions. It requires accepting a delay of several months between the work and its appraisal, which jars with the annual rhythm of HR processes. It requires, finally, distinguishing correctness from luck, which is possible only if the assumptions were written at the moment of the decision. Here we are back at the previous point.
Elsewhere, the same houses have been doing this for a long time and under obligation. The effectiveness check on a corrective action reopens, at a fixed deadline, a decision taken one or two years earlier and rules in writing on whether it produced the expected effect. The periodic benefit-risk report does the same on a product’s safety assumptions, and the annual quality review on process assumptions. These arrangements run, people go through the motions of them under obligation, and most have emptied out into a compliance ritual. That is the best argument I have. One gets box-ticking when the loop serves to prove that the work was done properly, and one gets something else when it serves, first, whoever feeds it.
But this evaluation has a property nothing else has. It measures what counts. And it has a considerable side effect, which is perhaps the best argument in its favour. The day an organisation evaluates its experts on the correctness of their arbitrations, it mechanically creates the incentive to become a better judge. Today the incentive bears on output, so people produce, and AI lets them produce more, which makes them better at what counts less and less.
One obtains exactly the behaviour one rewards. That is true everywhere and it is true here too. Convincing your experts to judge better will be of no use until your organisation has the courage to reward something other than what it sees.
X. What You Thought You Were Buying
There is a sentence one often hears, and it sums up this whole text better than I do, about an expert on the way out. We are going to have trouble replacing him.
Ask what will be hard to replace. They will never mention the syntheses. They will mention the time that person said one should not go there, and was right.
Everyone therefore already knows where the value is. Yet too few people have ever written down how much of it lies there.
That was bearable as long as the two halves of the trade came in the same body, at the same price, without their having to be distinguished. That is less and less the case today. One of the two has become available, abundant and cheap. The other has stayed exactly as rare as before, and it now alone carries the whole of the value, inside organisations that have neither named it, nor measured it, nor protected it.
I would like to finish by setting aside two misunderstandings, because they are the most frequent reactions to this reasoning and neither of them holds.
The first consists in concluding that fewer experts are needed. It is the opposite. An organisation producing a hundred times more hypotheses needs more people capable of deciding, not fewer. It simply no longer needs them to spend their time collecting. The confusion comes from continuing to count experts in units of output, when they should be counted in capacity to arbitrate.
The second consists in concluding that the trade will be impoverished. Here again I believe the contrary, and not out of optimism. A trade from which the thankless tasks are removed, keeping only the part that demands discernment, becomes harder, more demanding, and probably more interesting. The problem arises elsewhere. The trade becomes harder at the precise moment one stops training people in it.
It is therefore not a problem of artificial intelligence. It is a question that had been waiting a very long time to be asked: what exactly are you paying for when you pay an expert?
As long as the answer was not necessary, not having it cost hardly anything. Today, it has just become necessary.