An unattended wooden loom halted mid-pass in a dark mill hall, a broken thread hanging in the light from a single high window, unfinished cloth spread across the floor

LLM Jargonaut 101

A course listing for a profession that does not exist yet, written by someone who has been practising it for two months without a name for it.

LLM JARGONAUT 101. Integrating language models into a domain you do not understand. Three credits. Students learn to build closed verification loops around a producer whose competence cannot be inspected, in a field whose competence they do not possess. Topics: choosing a grading key by rung; validating the grader; deciding, by measurement rather than instinct, where a human belongs in the loop; writing scope and stop as executable conditions; leaving a trace a hostile reader can audit. Practicum: arrival in a foreign domain. Graduation by an exam the candidate can lose. Prerequisites: none in the domain, and a great deal in something else.

There is no such course. There is, increasingly, such a job, though it is rarely advertised under one title: someone walks into a bank knowing nothing about Know Your Customer, or into a newsroom knowing nothing about journalism, and is asked to make a language model useful there without breaking anything the institution cannot afford to break. The person is not a prompt engineer, because prompting is an input to the work and not the work. They are not a machine-learning engineer, because they never open the model. They are not a banker or a journalist, and never will be. What they know is a third thing. This piece is an attempt to write down what that thing is, using eight research literatures that already own most of its parts and, as far as I could find, mostly do not read one another.

One honest framing before the syllabus. The job is one seat of a two-seat pattern. The other seat is held by someone who knows the domain. The market already staffs both, the first under job titles like evaluation engineer; what it has not yet done is describe the first seat as a discipline, a body of knowledge a person can be shown not to have. This course is about that seat.

A warning before the argument. I am the author of the method this piece describes and the executive in the one body of evidence it offers. That evidence was gathered on my own coast, in a domain I can practise. The method is what I can claim to have run. The profession is the prediction.

Where the name comes from

The word is borrowed from Peter Watts's Blindsight, where a synthesist — jargonaut on the street — is the professional who stands between minds that have outgrown human comprehension and the institutions that still have to act on their output. He does not understand the physics or the biology. He "specializes in processing informational topologies": he preserves the shape of what the experts produce and trusts the meaning to survive the ride down to the people paying for it. He is, by his own description, "the curtain."

Watts builds him in order to take him apart. On my reading of the book, the synthesist ends it with no answer key, no defined scope, no reader with the standing to overturn him, and no signature under anything he transmits — and its most cutting line is that a man forbidden opinions on the job does not stop having them; he attributes them to the systems he observes. The novel is not a model for this profession. It is a portrait of the job done as pure transmission, and every element the book strips from him is an element the real job has to put back.

A frustrated colleague tells the synthesist that an alien species of "nothing but surfaces" should be his dream assignment, and he fails at it completely. I read that scene as a question: can an outsider who has only surfaces reproduce what an insider knows? That reading is mine, not Watts's. But a sociologist has run a nearby test.

Harry Collins spent three decades sitting in on gravitational-wave physics without ever doing any. In 2006 a physicist set seven technical questions; Collins and a real gravitational-wave physicist both answered; nine other physicists in the field judged the transcripts blind. Seven of the nine said they could not tell which was which. Two picked Collins as the physicist. The result is not that experts are fools. It is that there is a real, acquirable kind of expertise — Collins and Evans call it interactional — that lets someone speak a domain fluently, judge its arguments, and follow its meetings, without being able to practise it. Their term for using a neighbouring competence to make judgments about a field one cannot contribute to is referred expertise, and it is the nearest technical name I have found for what the person in the course listing has.

Collins is also the source of the warning, twice over. His fluency took decades, not the three months of meetings a deployment allows. And fluency fools experts: the jargonaut's own confidence in banking is the first thing they should distrust.

Why the course exists

Generation got cheap. The previous piece in this series was precise about what did not: comparison got cheap, judgment did not, and the cost curve collapsed only where outputs are short and closed. A recent economics preprint calls the gap between what can now be produced and what can be checked the measurability gap, and locates the scarce resource where this piece does — in the grader, not the generator. The profession lives in the region where that holds.

What the gap does inside an organisation is familiar to anyone who has managed people. I am the chief technology officer of a company whose lead developers know their projects far better than I do. I cannot compute what they will produce. I manage them anyway, with soft procedures — reviews, a definition of done, a person whose signature is on the release — and the procedures work well enough that nobody calls it a miracle. Management is an institution for acting on what you cannot compute, and the obvious move is to point it at the model.

The reason that fails is that the procedures rest on assumptions about the thing being managed, and the assumptions are about people. I count five, and the count is mine. A managed person has a persistent identity that learns from correction. Their competence is, if not contiguous, at least reported: they tell you, or show you, when a task is outside what they can do. They have some metacognition. They respond to incentive and reputation. And when they fail, they tend to fail gracefully — a worse version of the right thing rather than a fluent, confident, differently-shaped thing. People violate every one of these sometimes; Collins passing as a physicist is a human failing the last one on purpose. Language models violate each of them more often and less predictably, and the literatures that study delegation, oversight, and control have each noticed a piece of that from their own side.

So the management gap is not the ordinary one. It is management minus its human assumptions, and no discipline I could find owns the subtraction. The nearest candidate is bank model risk management, which for fifteen years has been the most codified practice for governing a system whose internals you cannot verify under a standing duty to leave a trace. It would be convenient to say the bank already has the chair. On 17 April 2026 the US banking agencies superseded their model-risk guidance with a shorter document that places generative and agentic AI "not within the scope of this guidance" in a footnote, mentions the technology nowhere else, and promises a request for information that had not been issued when I checked in September. Other supervisors disagree on the record about whether the old frameworks stretch to cover this; one European Central Bank board member has said they cannot. The chair is at best pending.

An objection I owe to the model-risk literature, though nobody there has put it quite this way: model validation never was management. It was built for non-agentic artefacts, and what language models do is re-import human-like properties — open-endedness, apparent judgment, non-reproducibility — into a discipline that had defined them out. On that reading the profession is validation plus something rather than management minus something. I think it sits on the seam, and this piece subtracts from one side and adds to the other without settling which description is primary.

The syllabus

What follows is the body of knowledge, compressed to seven moves. The seven are my compression of a longer curriculum, not a canonical list. Each move ends in something that executes — a count, a rate, a diff, an agreement score — which is a definition of what the profession does, not a finding about what a domain needs; the practicum is where the two are made to meet.

The producer is measured, not read. The first thing to learn about the model is that it cannot judge its own output, and that this is normal: a loom cannot judge cloth, a controller does not certify its own stability. One tradition in control engineering built its fault detectors and runtime monitors outside the plant, using only what the plant emits, because components do not announce their own failures. The jargonaut does the same. They measure calibration on a rationally subgrouped set — by input type, never pooled, because a pass rate over a mixed population is a number about nothing — and they find, at least once, an input type where the better-prompted run is worse, so that the jaggedness of the thing has been seen with their own eyes. Watts's synthesist read surfaces. This syllabus does not teach that. Calibration is measured, not read.

The key is chosen by rung, and the top rung is negotiated, not conceded. Sometimes the domain hands you an answer sheet: the style guide, the spelling of names, the sanctions list. Sometimes it hands you hard rules and soft ones: never name a minor victim; keep the house voice. Sometimes it hands you a procedure — the two-source rule, the hold-for-legal workflow — and what you grade is whether the procedure was followed, from an instrumented trace. And sometimes it says we know it when we see it. The assurance standards have a verdict on that last rung for their own purposes: "vague descriptions of expectations or judgments of an individual's experiences do not constitute suitable criteria" for an assurance engagement. The jargonaut's move is to sit with one named domain expert and fifty real outputs, open-code the failures, cluster them, count them, and write a rubric. Be honest about whose that rubric is. The verdicts are the expert's; the clusters are the coder's construction, and the coder is you. So the rubric goes back to the expert to keep or dismiss, row by row, and then to a second domain person cold, and agreement is measured against a threshold registered in advance. Where agreement never clears it, the written output is that this rung stays with the institution's judgment. The audit standards also say how much of the domain the assurer must hold, and it is less than practice: "a sufficient understanding of the field" to evaluate the expert's work, with the responsibility for the opinion not reduced by having used one.

The grader is separate, and it is graded. Four of the fields I read require the detector to be a different thing from the producer. The jargonaut's grader is usually a second model with a rubric, and the previous piece's warning applies with full force: a second model from the same family is the same instrument sampled again, not a different thing. Nine judges from seven families supplied roughly two independent votes' worth of information in one recent measurement, and error correlation rises with capability. So separation is bought two ways at once — a different family, and validation against the named human on a held-out set, class-conditionally, by true-positive and true-negative rate rather than accuracy, because the failures you care about are rare and a judge that says "pass" to everything scores well on accuracy. Then one more test: delete the rubric from the judge's prompt and re-run. If agreement with the human barely moves, either the judge was never reading the rubric or the rubric was redundant with what the judge already believed; removing one thing tells you about one thing, and either answer is worth having before you ship.

The loop runs with the prediction written first. State the expected result before the run. Re-run the controls per batch, because the rater changes underneath you between Tuesday and Thursday. Give the loop a dwell time: you cannot swap prompts and models faster than it settles and still know what a measurement means. And build one detector for the rare catastrophic output whose rate never moves the mean, because a chart of averages will not see it.

The human is only at the top of the ladder. The human-factors literature here is old and, with one dissent worth naming, consistent. Bainbridge's ironies: the operator is left to monitor the thing they can no longer do. Complacency in monitoring is not cured by expertise or practice. Better interfaces for reviewing agent traces made reviewers faster and more confident without making them more accurate. Individuals largely cannot perform the oversight that mandates assign them; the lever that works is institutional. The dissent is Dekker and Woods, who argue that any ladder allocating functions between people and machines rests on a substitution myth; I use a ladder anyway, and say so. The literature suggests the human-in-the-loop decision is a rate — high override rates, or overrides that systematically improve the output, may signal the model is the problem — and a budget of expert attention. The method grades every failure row by how little its detection depends on a person: eliminated · designed out · a test · a gate or probe · a visible state · a surname · none. The ladder was derived on software and its transfer is untested. The rule beneath it is a design principle I hold rather than a result I can cite. A surname is a rung, not a solution. Inside scope, the budget buys sampled review at human pace, on the loop's own outputs, at a rate the loop can afford; that is a surname rung and it is defensible. What it cannot buy is a person reading the machine's selection at the machine's pace, catching rare, fluent errors as they stream past. That is Doctorow's reverse centaur, and a loop built that way does not get my signature. The method has no more authority than that, and says so below.

Scope is closed outputs; stop is a cord in the operator's hand. Safety engineering has the right concept for scope — the operational design domain, a written boundary of demonstrated competence plus a detector for leaving it — and the honest admission that nobody has a coordinate system in which to draw that boundary for a language model. That remains open. The working heuristic, taken from one software study and not yet tested elsewhere, is that scope is wherever the domain has closed outputs — short-horizon and comparable. Where the domain does not have them, the key move can sometimes build them, and building them is a choice the loop makes about where to work, not a boundary it discovers. For stop, the oldest artefact in the syllabus is the best. Sakichi Toyoda's loom halted itself on a broken thread. It did not know what good cloth was; it knew one detectable abnormality and was wired to stop on it. The andon cord extended that authority to the operator at the station — the person doing the work, not the expert who is never in the room. The jargonaut writes the design domain in one page, builds the exit detector beside the model using none of the model's internal state, and hands the cord to whoever runs the loop day to day. Then a week is run and every pull is classified as right, wrong, or late against the key the expert signed before the week began. The classification grades the detector and the key, never the hand on the cord: a late pull is a debt on the design, held by one of the two signatures below, not a mark against the operator.

The trace is instrumented, and two people sign it. A model's account of its own run is not a trace. The trace is the record the loop writes about itself from outside, and the audit tradition supplies the distinctions that matter: global versus local explainability — a traceable pipeline satisfies a prudential supervisor and does not satisfy an individual who is owed the reasons for a decision about them — and limited versus reasonable assurance, where naming which one you are selling is most of the ethics. The trace manufactures the accountable party, and the model cannot be that party. Neither can a name that cannot evaluate the domain, on its own; that would be the accountability sink the previous move forbids. So two names go on it. The domain expert signs the key — the verdicts were theirs and the rubric built from them went back to them row by row. The jargonaut signs the process: that the grader was validated, the controls re-run, the scope written, the cord held. Where a failure row sits on a surname or on none, the trace carries a debt row — what is silent, the check that would catch it, where it would run, who holds it until then, since when — and the holder is the jargonaut, because a debt is a gap in the process. One failure has no signature under it, and the trace says so: the fluent output that passes because the key does not cover it. That gap is between the key and the domain. It is the top rung, it belongs to the institution's judgment, and no name on the trace claims it.

Here is one of these moves when it is not a paragraph. The evidence is my own company, and it comes with the warning from the top attached: the domain is software, and I can practise it, so this run says nothing about the knowledge gap. What it says something about is the loop. The same real requirement was given to several arms and graded clause by clause against a rubric written before the outputs existed, with a fault register of forty-three findings that an independent reader extended by eight and overturned by four. On the first pass the method's output carried none of the five clauses of what was asked. After tracing every fault to the rule that produced it and repairing each rule with a detector attached, the second pass carried all five. Six cells of the first register were wrong, all six in the direction that flattered the finding; they are marked in the record, not fixed. No ticket from that run has been executed, so the documented experiment has evidence and no proof. The lead's know it when we see it became a rubric a second person could apply — and the inter-rater agreement this syllabus demands was not measured. That is the standard I am holding others to, not yet met by me.

Arrival

The practicum is the only place domain content enters, and it enters in an order.

Acquire the language first — sit in the meetings until you can follow them — and treat your fluency as a warning. Name the one person whose pass or fail is ground truth, and get in writing what their signature covers. Then elicit the loss list: what counts as an unacceptable outcome here — a libel, a burned source, a sanctions miss. The safety tradition calls this hazard analysis and is precise about who supplies the content. You extract; you do not invent. The losses and the verdicts are the domain's. What you build from them — the clusters, the rubric — goes back to the domain to keep or dismiss before it is anyone's key. Then sound a charted channel: two cases the domain already holds firm and opposite opinions about. If the apparatus cannot separate them, stop. Nothing after that would be a number. This is the one result in the course that counts against the profession rather than against a citation, and it is written down before the run.

Then the cheapest test in the kit. Same input, three independent runs, no shared context, read side by side clause by clause. When we did this to a one-line brief about exporting a user's data, three fluent specifications differed on a single word nobody had defined — archived — and that undefined word was the defect that had already shipped. The outsider's contribution is exact and limited: three runs will differ on many words, and the outsider can list them all without knowing the domain, but cannot rank them. The insider can rank them and reads straight past, supplying the missing definition silently. Neither finds archived alone. This is one worked case, in my own domain, and the mechanism is my inference. Adjudicate with the expert. Never merge.

Plant a contradiction and require a refusal. Delete a section and see whether the output moves. Run the napkin. Then build, and leave. The instruments travel. The map stays.

Three objections belong here, and two of them have no ready reply.

The first is a field experiment. Six hundred and forty Kenyan entrepreneurs were given a language-model business adviser over a messaging app. There was no average effect. Strong performers gained on the order of fifteen percent; weak performers lost about ten. The authors attribute the gap not to the advice given but to which advice the entrepreneurs chose to implement. My reply is that this was an unmediated channel — no key, no scope, no grader, no stop — which is what the loop replaces. But the finding has a second edge that the reply does not blunt. The loop inherits the calibration of whoever writes its key, and this is evidence that the person closest to a domain is not always the person who can judge advice about it. If the named expert is one of the losers, every executable number downstream is precise nonsense, and the loop's only domain oracle is that person. I have no design that detects this from inside.

The second is philosophical. Neil Levy has argued that a novice who applies the criteria of expertise to an expert domain from outside is trespassing, however carefully they do it. Referred expertise is the licence this essay claims. Where the negotiation over the top rung fails and the rung stays with the institution, I concede the licence is void there. Where the negotiation succeeds and a rubric comes out of it, Levy's objection is that the negotiation itself was the trespass — the outsider decided which of the expert's verdicts clustered together. To that I have no answer.

The third is that the profession may already exist under another name. Evaluation engineers already pair with a named domain expert and already run most of this syllabus. If so, the course is a renaming, and what is new in it is the claim that the first seat is a discipline: a body of knowledge with a syllabus, an exam, and a way to fail. That claim is conceptual, not empirical; no number of companies would settle it. I think it holds.

Lineage

The profession has ancestors, and they are not all honourable.

Frederick Taylor could not handle pig iron. He built the measurement, and a profession, on the claim that "the science of handling pig iron is so great and amounts to so much that it is impossible for the man who is best suited to this type of work to understand the principles of this science, or even to work in accordance with these principles without the aid of a man better educated than he is." That is prerequisites: none in the domain, stated in 1911 with the same self-serving edge. Fayol added that the managerial ability is a different ability and should be taught. Braverman's reading of what they built is the darker one: extract tacit knowledge from those who have it, encode it as procedure, hold the monopoly. The rubric built from a senior editor's fifty verdicts is the editor's judgment, clustered by an outsider and made executable, and nothing the loop measures says whether the editor ends up better or worse off for it.

Cory Doctorow gives the ethics a shape. A centaur is a human head on a machine body: the person directs at human pace and the machine assists where it is reliably good. A reverse centaur is the machine's head on a human body: the machine sets the pace, the human catches its errors at that pace, and absorbs the blame when it fails — the accountability sink, in Dan Davies's phrase, or Doctorow's own "crumple-zone for a robot." The moves above can be assembled into either. This method is not neutral between them, in the way a building code is not neutral about fire exits: it has a rule, and the rule's only enforcement is that the jargonaut does not sign a loop that breaks it. A withheld signature is not an incentive, and reverse centaurs are built on purpose because they are cheaper. The syllabus can teach the difference and name the choice. It cannot make it for you.

Watts's synthesist suspected his own guild was "some grand consensual con." The honest version of this profession keeps that doubt on the syllabus.

The exam

Refutation is cheap in software, where dead machinery is found by deletion and diff. In a foreign domain it costs one thing extra: knowing what a gate was supposed to catch, which is the loss list, which the domain supplies. With that in hand the asymmetry the previous piece described carries over. The loop can tell the newsroom what is dead — the gate that never fired, the step whose removal changed nothing, the rule everyone followed and nobody had written down. Whether the newsroom is good is an endorsement, and endorsements are first-party unless someone independent gives them. The jargonaut built the process, so their own word on it is limited assurance on their own work, and it should be sold as that. Reasonable assurance on the process belongs to the independent reader with the standing to overturn. On the domain, only the domain's signature speaks.

So the graduation exam is built to be lost. A domain the candidate has never worked in. A domain expert who supplies the verdicts, signs the key, and is not the examiner. A rubric registered before the run. A napkin — one senior domain person, one day, no apparatus — scored on the same substrate and published beside the candidate's number. An independent reader with the standing to overturn. The napkin is the exam's honesty device, and it has form: the one time we ran it against our own method it did not obviously lose, and we had written down in advance that in that case we would claim nothing. Whether a domain senior on a napkin is the right baseline in a foreign domain is still open; the exam assumes it and cannot test it.

The candidate passes on three conditions, each of which executes.

  • The loop found at least one dead thing the domain did not know was dead.
  • No failure row that the loop ships with puts a person at machine pace, and every row on a surname carries a debt row naming the holder and the date.
  • The trace, handed to the reader, supports every claim in the candidate's summary. One unsupported claim fails the exam.

The candidate does not pass by beating the napkin. The napkin is allowed to win. A candidate who reports that it won, with the number beside their own and the trace under both, has demonstrated the only thing the profession can honestly promise: a loop around a producer nobody can inspect, in a field the builder cannot practise, that measures what it can close and writes down what it cannot.

The synthesist ended his voyage with one order: survive and bear witness. The witness in this course signs, and so does the domain.

Prerequisites: none in the domain.

No comments yet