If you had to guess who uses AI the most inside a company, most people would say the engineers.
On August 12, OpenAI’s Chief Economist Aaron Chatterji, David Holtz, and coauthors posted a working paper, “How Organizations Use AI: Evidence from ChatGPT” (official PDF), that links ChatGPT Enterprise usage records to employee job titles, task classifications, and public-company financials. The guess turns out to be half right. By headcount, engineers are the largest group of active users. But by messages per user, the people sending the most within a given company are entry-level employees and interns (seniority here is inferred from job titles; the paper has no hire dates or tenure data), who send roughly eight to nine more messages per week than the average active user at their firm. Executives sit below average, and the more senior the role, the fewer the messages.
I went through the paper and checked the key numbers against the original. Data caveats first, then three patterns I think carry real information.
Whose data this is
The sample is organizations that adopted ChatGPT Enterprise between January 2024 and March 2026. The employee-level analysis covers 1,764 organizations and 17.4 million messages. The task-classification sample is 973 organizations and 8.7 million messages sorted into 60 task categories, but classification data only begins on October 30, 2025, so it can’t reach further back.
Two things belong on the table up front. First, this is OpenAI analyzing its own product with its own data, authored by its own chief economist. The interest is obvious. Second, the paper only sees ChatGPT Enterprise usage. It cannot see companies using Claude, Gemini, in-house applications, or employees’ personal accounts. The paper concedes this itself: “non-adopters” are simply firms that couldn’t be matched to an Enterprise account, not firms that don’t use AI. Every adopter-versus-non-adopter comparison below reads more accurately as “what OpenAI’s enterprise customers look like” than as “what companies using AI look like.”
Pattern one: half the growth comes from existing customers using it more
Between June 2025 and March 2026, aggregate output tokens (the text the model generates in reply, a rough gauge of usage) consumed by ChatGPT Enterprise customers grew sevenfold. Broken out, firms that had already adopted before June 2025 grew their usage roughly fourfold over the same window. About half of total growth happened inside firms that had already adopted, rather than coming from new customers.
That split matters more than the sevenfold headline. If growth came entirely from new customers, it might just be a sales achievement. Half of it coming from the existing base means the companies that bought in first are, in aggregate, still ramping up. One caution: the paper gives the aggregate curve for that cohort, not per-customer breakdowns. Whether usage is deepening broadly or a few firms are surging while other seats sit idle is not something this data can show.
The token metric itself deserves a question mark. A sevenfold rise in output tokens does not translate to “people used it seven times more.” Three forces could plausibly feed the total: more people sending more messages, longer replies per message (reasoning models spread widely during this period, which could mean more generated text for the same question), and multi-step tool calls behind a single task. That three-way decomposition is my conjecture, not the paper’s: it does not publish a message-count growth curve on the same basis, nor a per-message token trend, so the data can’t apportion the three, and I can’t verify that each force actually contributed. The paper does clarify one piece: enterprise token output overwhelmingly comes from ChatGPT itself and non-agentic tools, with agentic products like Codex still a small share. That caps the contribution from agentic products, though it says nothing about tool calls inside ChatGPT itself, and “the model got wordier” can’t be ruled out either. And token consumption is not output value. The paper says plainly that it measures use, not productivity.
Pattern two: adoption concentrates in firms that were already strong
Among adopters that could be matched to US public companies, the median 2024 adopter had revenue of $2.275 billion, against $210 million for non-adopters, a gap of more than tenfold. Market value, headcount, and R&D spending show similarly wide gaps.
Bigger firms adopting first is unsurprising: firm size itself is strongly associated with adoption in every specification the paper runs. The interesting part is which variables still predict adoption after controlling for size: R&D stock per employee and capitalized software are both positive, and the strongest association among these is SG&A stock per employee. SG&A is the selling, general, and administrative line on an income statement, covering sales operations, administration, and management overhead. The paper accumulates past SG&A spending into a capital stock, depreciating it at 20% a year, and uses that stock as a proxy for organizational capital, meaning a firm’s built-up intangible capabilities in process, management, and sales networks. In the other direction, physical assets like plant and equipment are negatively associated with adoption once size is controlled for.
This lines up with an old result from research on the last IT revolution: Brynjolfsson and Hitt argued in 2000 that the value created by IT investment depends heavily on whether complementary organizational investment keeps pace. If AI follows the same rule, then in the short run it amplifies incumbent advantages: firms that already have organizational capability adopt first. Whether they also use it most intensively is a separate question, and the paper’s answer cuts the other way: conditional on adoption, measured use per employee is lower at larger firms. A reminder is due here. This is correlation, not causation, and strong firms buying first is itself a selection effect.
Pattern three: usage intensity inside firms runs upside down
On who uses it within a company, the paper offers two measures. By headcount composition of active users, engineering and technical roles are about 11%, the largest functional category, and the manager-and-director seniority band is about 24%. By messages per user, the ranking flips: entry-level employees and interns send roughly eight to nine more weekly messages than the average active user at their company; analysts and marketing and communications roles also sit above average; executives, founders, and partners sit below. The more senior, the fewer messages, in a clean gradient.
The measure needs to be stated precisely. The paper counts messages sent per active user per week. It does not count tasks initiated, and it does not separate the number of conversations from the length of each one. The authors themselves caution that message volume measures intensity of use, not economic importance. That leaves a reading the data can’t exclude: senior employees handle fewer, larger, harder tasks, with much of the work happening away from the keyboard in thinking and design, while junior employees hold many small units of output, each of which converts directly into a few prompts. More messages might mean more tasks, or just finer-grained tasks. There is also the nature of the job itself: senior people spend much of the day in meetings, coordination, and sign-offs, activities that simply don’t generate many tasks you can hand to ChatGPT. Fewer messages need not mean reluctance. It may just mean less promptable work on hand.
The task content backs this up. Executives’ messages skew toward orientation tasks: topic overviews, facts and figures, legal, regulatory, and tax questions. Junior and front-line employees’ messages show up disproportionately in the common production-oriented task categories, the kind of work that turns directly into drafts and deliverables. ChatGPT today is a production tool, and whoever is doing front-line production has prompts to send.
Put this finding next to another thread and the picture gets more complicated. The Stanford Digital Economy Lab’s “Canaries in the Coal Mine,” using ADP payroll data, finds that since generative AI took off, employment of 22-to-25-year-olds in the most AI-exposed occupations now stands about 19% below where it would be had it kept pace with similarly aged workers in less-exposed occupations. By that same measure, the gap was 15% at the July 2025 data vintage; the August 12, 2026 revision extends the data through June 2026 and puts it at 19%. (The 13% figure in the initial August 2025 version was a regression estimate adjusting for firm-level shocks, a different measure, so the two numbers aren’t directly comparable.) The gap operates mainly through hiring: companies are hiring fewer young workers, not laying off the ones they have. The group using the tool hardest is the same group whose entry-level openings are shrinking fastest. Two readings are available. The optimistic one is that AI pushes out the capability frontier for junior workers: analysis that used to require years of accumulated experience, code they couldn’t have written alone, is now within reach through the model, converting a slice of “experience” into a tool available on demand, so genuinely strong juniors will overtake senior colleagues faster than before. The pessimistic one is that junior work is the most AI-substitutable, which is why juniors are first to be told to use it heavily and first to become optional hires. The OpenAI paper notes this connection itself, but usage data can’t decide between the two readings, and I can’t either. Both mechanisms can be true at once.
A few operational judgments
For people driving AI rollouts inside companies: the paper doesn’t compare rollout strategies, and it can’t see how usage actually spreads through an organization, so what follows is my judgment from the descriptive structure, not a tested recommendation. The heavy senders are front-line production roles and early-career employees, and executives use it lightly. Designing the rollout around executive use cases and leadership-keynote launch events runs against that structure. The real landing zone is tasks that cut across every department: more than half of active users have done documentation or technical writing, and across industries message volume concentrates in the same core categories of documentation, technical work, and communication, with industry differences showing up mainly at the margins of who touches which tasks rather than where messages concentrate. Roll out the cross-cutting tasks first; role-specific deep use cases are the second layer.
For people watching the industry: firm size predicts adoption strongly, and once size is controlled for, the variable most strongly tied to adoption is preexisting organizational capital. The “AI levels the field between big and small companies” narrative has not shown up in enterprise adoption, at least not in this data: the firms adopting first are the ones that were already strong, even if per-employee usage after adoption runs lower at the largest of them. Whether this goes on to widen the gap between firms is something the paper doesn’t measure, only flags as possible. And keep the single-product lens in mind: smaller firms may well be using cheaper alternatives.
For individuals, especially early-career ones: heavy usage and entry-level contraction are happening at the same time. The tool won’t choose which mechanism comes true for you. What you can do is steer your own use toward building experience and moving up a level, rather than stopping at passively filling an output quota. That is easy to say, and I admit there is no standard playbook for doing it.
The number worth remembering from this report is not the sevenfold growth. It is the half of that growth coming from existing customers, together with SG&A’s predictive power for adoption. Both point the same way: whether a company can actually put AI to work is tightly bound to the organizational capability it already has. Whether this round of returns depends on complementary organizational change the way the last IT revolution did is beyond what this data can test. What it can confirm is that the firms off to a fast start are the ones with thick organizational capital.
References
- How Organizations Use AI: Evidence from ChatGPT (arXiv:2608.12236) — all core data: sample composition, sevenfold/fourfold growth, the note that tokens come mainly from non-agentic tools, adopter vs. non-adopter financials, SG&A stock construction and regression results, usage intensity by function and seniority, task distribution, stated limitations
- Official PDF (OpenAI CDN) — the official release of the same paper
- Dr. Ronnie Chatterji named OpenAI’s first Chief Economist (OpenAI) — official source for Chatterji’s Chief Economist role
- Brynjolfsson & Hitt, “Beyond Computation” (JEP 2000) — the classic result on complementarity between IT investment and organizational capital, used here for mechanism comparison
- Canaries in the Coal Mine? (Stanford Digital Economy Lab) — source for the ~19% kept-pace employment shortfall of early-career workers in AI-exposed occupations (August 12, 2026 revision, data through June 2026; the same measure stood at 15% at the July 2025 data vintage)