Analysis & Insight

Producing the analysis is the part that automates

Routine extraction, querying and charting have become fluent. Ambiguous and novel analysis has not. The model sounds the same either way, which is the whole problem — and, unusually for this site, the demand for these occupations is projected to rise rather than fall.

Value moves from producing the analysis to certifying it: knowing where the model fails, owning what the numbers mean, and catching output that is confidently wrong before it reaches a decision.

Measured

The 3 occupations this covers

AI performs unevenly across a jagged frontier: fluent on routine extraction and charting, unreliable on ambiguous or novel analysis — and equally confident in both cases. Value moves from producing the analysis to certifying it.
Occupations in the Analysis & Insight cluster with their measured AI exposure percentile and the observed automation share of AI usage.
OccupationMeasured exposureObserved automation share
Operations research analysts95th percentile53% — too close to call
Management analysts93rd percentile45% — too close to call
Computer systems analysts92nd percentile70%

The automation share is measured from observed AI usage, published by Anthropic under CC BY 4.0 — a vendor reporting on its own product, and worth reading as such. It describes how people use AI for this work, not how much of the work AI can do. Exposure percentiles are our composite of the measured sources set out in the methodology.

  • Confident wrong answers that pass casual review and reach decisions
  • Semantic and definitional drift — the same metric meaning different things per tool
  • Self-serve AI analytics bypassing the analyst until something breaks
  • Judging your own reliance on a tool that is right most, but not all, of the time

The shift

The measurements and the projections disagree

This page argues something different from the others on this site, because the evidence is different. These three occupations are among the most AI-exposed we measure — and every one of them is projected to grow much faster than average. Both are true, and the reason they are both true is the argument.
  • The largest of the three occupations is projected to grow strongly over the decade, and the statistical agency's entry for it does not mention artificial intelligence at all. That absence is itself a finding.

    Management analysts: 1,077,100 jobs, projected to grow 10% from 2025 to 2035 — “much faster than the average for all occupations” — with about 94,100 openings a year. Median pay $101,860.

    US Bureau of Labor Statistics, Occupational Outlook Handbook · 2026-08-27 · verified 2026-09-07

  • For the systems and process analysts in this cluster, the agency names AI explicitly — as a reason to hire, not a reason to cut.

    Computer systems analysts: 544,400 jobs, projected +8% to 2035. As organisations expand IT, “including artificial intelligence (AI), computer systems analysts will be hired” to design and install new systems.

    US Bureau of Labor Statistics, Occupational Outlook Handbook · 2026-08-27 · verified 2026-09-07

  • And the fastest-growing of the three is the one whose work is most obviously quantitative. Better analytical software is given as a reason the field expands, because it widens what the analysis can be applied to.

    Operations research analysts: 113,100 jobs, projected +12% to 2035, because “improvements in analytical software have made operations research more affordable” and applicable more widely.

    US Bureau of Labor Statistics, Occupational Outlook Handbook · 2026-08-27 · verified 2026-09-07

  • The finding that reconciles high exposure with rising demand, from administrative payroll records rather than a survey. Exposure does not predict decline; what the AI is used FOR does. Where it stands in for the person, employment falls. Where it works alongside them, it does not.

    “Declines are concentrated in occupations where AI usage primarily substitutes for human tasks”; and “where usage primarily complements workers, employment is flat or rising”.

    The authors are explicit that these are “early, descriptive indicators — canaries in the coal mine — rather than causal estimates”.

    Brynjolfsson, Chandar and Chen, Stanford Digital Economy Lab, ADP payroll data through June 2026 · 2026-08-12 · verified 2026-09-07

  • The same work finds the adjustment is not falling on everyone equally. It runs through hiring rather than firing, and it lands on the entry point — which is where analysis has traditionally been learned.

    Employment of workers aged 22–25 in AI-exposed occupations now stands “19% below where it would be had it kept pace” with less-exposed peers. Experienced workers show no comparable gap.

    Brynjolfsson, Chandar and Chen, Stanford Digital Economy Lab · 2026-08-12 · verified 2026-09-07

The frontier

It is jagged, and it does not announce where it ends

  • The best-known experiment on this found large gains on tasks inside the model's competence — and, on a task designed to sit just outside it, that having the AI made trained professionals substantially worse than working without it.

    Across 758 consultants, those with AI access on the outside-the-frontier task “were 19 percentage points less likely to produce a correct recommendation”. The control group was right 84.5% of the time.

    The fieldwork was run in 2023 on GPT-4 and the frontier has moved since. Cited via a co-author's own retrospective because the publisher's pages are not openly fetchable.

    Dell'Acqua, Mollick, Lakhani and colleagues, Harvard/BCG field experiment, via co-author Karim Lakhani · 2026-03-16 · verified 2026-09-07

  • The reason a checking habit is not sufficient protection: when professionals pushed back on the wrong answer, the system did not concede. It argued, using the form of analysis rather than the substance.

    Challenged, the model deployed “structured reasoning and comparisons to make its flawed recommendation appear analytically grounded”.

    Karim Lakhani, co-author, Harvard Business School · 2026-03-16 · verified 2026-09-07

  • And the reason this is a standing job rather than a one-off training exercise. The boundary of what the model does well moves with every release, so no fixed review checklist stays correct.

    Organisations “cannot build static processes around a dynamic capability boundary”.

    Karim Lakhani, Harvard Business School · 2026-03-16 · verified 2026-09-07

  • Experts do not catch wrong output reliably, and the pattern of who fails is the opposite of comfortable. In a randomised experiment, the same incorrect judgement was accepted more often when it was labelled as coming from a machine.

    With 1,339 participants, an unduly harsh incorrect recommendation labelled AI-generated “increases the fairness gap by 0.302 points (P<0.01)” — a 22% relative increase over the identical human version.

    The task is grading, not business analysis, so read it as evidence about deference to labelled-AI judgement rather than about analytical review. Note who deferred most: the younger, more highly educated and more technologically confident.

    Goulas, Megalokonomou and Sotirakopoulos, PNAS Nexus, randomised labelling experiment · 2026-06-09 · verified 2026-09-07

  • The strongest evidence that certification is scarce is not an accuracy gap. It is that the people who wrote the answer keys for the industry's own benchmarks got the analytical question wrong more than half the time — on questions with a single correct answer.

    Expert re-analysis found “BIRD Mini-Dev and Spider 2.0-Snow have error rates of 52.8% and 62.8%”, with rank changes on the leaderboards ranging from −9 to +9 positions.

    This cuts both ways and we are citing it against ourselves too: it means published text-to-SQL accuracy figures, including the ones in the section below, rest on soft ground.

    Jin, Choi, Zhu and Kang, University of Illinois Urbana-Champaign, in Proceedings of the VLDB Endowment · 2026-01-13 · verified 2026-09-07

On the record

Someone has said this part out loud

  • A chief executive has publicly divided his company into people who build, people who sell, and people who measure — and cut the third group while revenue was at a record. Read the definition before deciding whether it describes you.

    “The vast majority of those we laid off last week were measurers.” — Matthew Prince, CEO, Cloudflare, on a cut of more than 1,100 roles.

    He defined measurers as middle management, finance, legal, internal auditing and revenue recognition — not business intelligence analysts. Stretching it to mean analysts would be exactly the overreach this page is trying to avoid. What it does show is that the category is being drawn, and that measuring is on the wrong side of the line.

    Matthew Prince, CEO, Cloudflare, reported by Jake Angelo, Fortune · 2026-05-21 · verified 2026-09-07

  • The company's own account of the reasoning is worth reading, because it is an argument about what measurement is for rather than about cost.

    AI systems, he argued, “can now measure an organization with a level of objective detail and precision” previously impossible. — Matthew Prince

    Matthew Prince, CEO, Cloudflare, via Fortune · 2026-05-21 · verified 2026-09-07

Practitioners

What the people doing this work say the hard part is

  • The structural reason analysis resists automation in a way software does not, from someone who spent a decade building tools for it: there is no test that tells you the answer was right.

    “Analysis isn't testable.” You find out whether the recommendation was good only after it has played out.

    A vendor founder — Mode was acquired by ThoughtSpot — writing about the category his products sit in.

    Benn Stancil, founder of Mode Analytics · 2026-02-13 · verified 2026-09-07

  • And the practical form the work takes now, described by someone who does it: not producing the number, but checking it against something authoritative, because the definition is where the ambiguity lives.

    “my next action is to cross check against some source of truth”, because “people don't know exactly what they mean by a metric”.

    Benn Stancil, “The context layer” · 2025-08-29 · verified 2026-09-07

  • The reliability paradox that makes this a permanent job rather than a transitional one: the better the system gets, the less closely anyone watches it, and the more expensive the errors that do get through become.

    “AI capability ≠ AI safety.” — Cassie Kozyrkov

    The body of the piece is paywalled; this is the headline formulation and the framing of her argument, not a reading of the full text.

    Cassie Kozyrkov, former Chief Decision Scientist, Google · 2025-06-10 · verified 2026-09-07

Against this reading

The case that certification is not the answer

The strongest objections here are not that analysts are safe. They are that the tools got good quickly, that there may simply be less work, and that the destination is a set of duties rather than a market.
  • The tools got good quickly, and any argument resting on their limitations dates fast. On the hardest public benchmark for enterprise querying, systems went from failing almost everything to topping it inside about two years.

    On Spider 2.0-Snow, frontier-model baselines of roughly 13–24% in 2024–25 have been succeeded by leaderboard entries above 90%.

    The top entries are self-submitted by vendors, on a benchmark whose annotations were found to be 62.8% wrong. Report them; do not treat them as measurements. The honest reading is that nobody has a trustworthy number in either direction.

    Spider 2.0 public leaderboard, Lei et al., ICLR 2025 · 2026 · verified 2026-09-07

  • If routine query production compresses far enough, the problem is not that the work changes character — it is that there is less of it. A named company shipped this at scale and published what happened.

    Uber's internal text-to-SQL tool cut query authoring from about 10 minutes to 3, with “about 78% saying that the generated queries have reduced the amount of time”.

    Dated 2024, and a company describing its own tool. Its limitations section is the counter-counter: it reports queries generated against tables and columns that do not exist, evaluation that cannot cover the space of business questions, and identical runs producing different outcomes.

    Uber engineering blog, on QueryGPT · 2024-09-19 · verified 2026-09-07

  • The sharpest objection to a certification premium is that AI narrows the gap it would be paid for. Where this has been measured, the least experienced gained most and the most experienced gained least.

    Across 5,179 support agents, AI assistance raised issues resolved per hour 14% on average and about 34% for novices, with near-zero effect for the most experienced.

    Customer support, not analysis, and fieldwork from 2020–21. It is here because its shape is the threat: a technology that compresses the novice–expert gap is a technology that erodes what expertise commands.

    Brynjolfsson, Li and Raymond, Quarterly Journal of Economics · 2025 · verified 2026-09-07

  • And the same experiment this page leans on for its central claim found large gains everywhere the task sat inside the model's competence. Most analytical work is inside it.

    Inside the frontier, consultants with AI completed 12.2% more tasks, 25.1% faster, and at significantly higher quality.

    Dell'Acqua, Mollick, Lakhani and colleagues, Harvard/BCG field experiment · 2026-03-16 · verified 2026-09-07

Where the work goes

The job on the other side of this

  • The destination is not a new job title. It is a set of duties being written into an existing one, at an ordinary company rather than an AI lab — and the word it uses for the central task is not produce.

    Axios, Analytics Engineer, $130,000–$155,000: “define, version, and maintain business metrics and semantic models”, and “review and certify dashboards for accuracy, consistency, and adherence to standards”.

    One posting at one company. Note what it implies about the market though: the certification work is being absorbed into an existing role rather than creating a new premium one.

    Axios job posting · 2026 · verified 2026-09-07

  • The same work exists in its most valued form at the labs building the models, and the job description is structurally what an analyst already does — define the measure, own the dashboard, make regressions impossible to miss, make it legible to whoever decides.

    Anthropic, Research Engineer, Model Evaluations: converting ambiguous notions of intelligence “into clear, defensible metrics that researchers, leadership, and the public can rely on”.

    Read honestly, this is a research-engineering role at a frontier lab requiring machine-learning infrastructure skills. It is not a job to apply for tomorrow, and it is shown here for the shape of the work rather than as a destination. Outside labs like this one we found no evidence of a real hiring market for evaluation as a standalone title — the duties are real, the category is not yet.

    Anthropic job posting · 2026 · verified 2026-09-07

We are not going to name a job title here. The obvious candidates are either a different profession wearing a familiar name, or freelance task work dressed up as a career step, and pointing you at one of them would undo the argument above. What changes is what you are accountable for; where that is worth the most is not yet a title we can stand behind.

Where does your own work actually stand?

The measurements above describe occupations. They do not describe you. The audit reads your own evidence — what you have owned, and what of it is repeatable somewhere else.

Get Started