A new paper on scientific agents system prompts reports that detailed, profession-specific instructions raised token use and estimated cost per response without a consistent accuracy gain. The authors released Scientific Agents, an open-source corpus of 503 such profiles, and tested them against four controls.

Each entry is a profession-specific AGENTS.md profile. The authors evaluated them with Gemini 3.8 Flash, accessed through OpenRouter, in the Pi agent harness. For teams that write long role-based instructions on the assumption that more detail helps, the finding raises a plain cost question, though it rests on one model and one study.

What did the Scientific Agents study test?

It tested whether a detailed, profession-specific instruction file makes an agent better at scientific work than far shorter alternatives. A system prompt is the standing instruction a model receives before the user’s request, and role prompting means telling the model to act as a described professional. AGENTS.md is generally described as a plain-text file that agents read for standing instructions; that definition is our background, not something the abstract spells out.

The authors’ claim has two parts. Detailed profiles raised token use and estimated cost per response, and they did not produce a consistent accuracy gain. The abstract gives no figures for either, and it names the Pi harness without describing it, so we do not describe it either.

What does each of the four controls isolate?

Each control strips a different layer from a matched profile, so each supports a different conclusion.

ControlWhat it keepsWhat the comparison can show
Minimal baseline (‘You are a helpful assistant’)Almost nothing beyond a generic roleWhether profile content beats a bare prompt
Opening role sentenceOnly the profile’s first sentence naming the roleWhether the long body adds anything beyond stating the role
Generic scientific rigor guideGeneral scientific care, no professionWhether generic rigor advice does the same job
Unrelated-domain profileA profile written for a different domainWhether matching the profession matters, as opposed to having any profile

We read these as a ladder of increasingly fair comparisons. Beating a one-line prompt is the easy claim. Beating an unrelated-domain profile is the hard one, because it asks whether the match to the profession matters or whether any detailed profile would do. That last comparison decides whether professional specificity earns its length, and the abstract does not say how it came out.

Diagram of one matched profile compared against four controls, from a one-line baseline to an unrelated-domain profile
The study’s design as the paper describes it: each matched profile is compared with four controls. The paper’s abstract does not report which performed best.

How does a longer system prompt turn into cost?

A longer system prompt costs more because its tokens count as input on every call that includes it. Under token-based pricing, instructions are not free background. An agent that resends its instructions on each call pays that overhead repeatedly, not once, which is why prompt length matters more for agents than for single-shot chat.

The paper’s figure is an estimated cost, not a billed amount, and the abstract does not size it. BriefFlash has covered OpenRouter, the gateway used in the study, which routes requests to many models and which Stripe will reportedly acquire.

Does ‘no consistent accuracy gain’ mean role prompts do not help?

No. ‘No consistent gain’ is weaker than ‘no gain’. It leaves room for some profiles or tasks to have improved while others stayed flat or got worse, and the abstract does not say which happened.

The study therefore does not show that role prompts never help. It shows, in the authors’ account, that any benefit was not reliable enough to count on. For a team working in one field, a pattern across many professions may say little about the one profile it cares about.

What the study leaves unanswered about scientific agents system prompts

We worked from the paper’s abstract, not its full text, so these questions stay open:

  • How much token use and estimated cost rose.
  • What the scientific tasks were, how accuracy was scored and how many tasks ran.
  • How the matched profiles, the rigor guide and the unrelated-domain profile compared with one another.
  • Whether profile length, content or both drove the cost.
  • Whether other models, or tasks outside science, behave the same way.

Gemini 3.8 Flash is the only model named. Our coverage of Gemini 3.8 models on AI Gateway concerns Live models, which are not the Flash model used here.

What should teams writing agent instructions take from it?

Treat a long role profile as a hypothesis to test, not a default. We think the study gives teams a reason to measure, not a reason to delete: it reports a cost increase alongside an accuracy benefit that was not consistent. Our confidence is moderate, since one model and one abstract carry the whole result.

A practical test copies the paper’s ladder. Run your own tasks with a one-line baseline, your profile’s opening role sentence and the full profile, and record accuracy and tokens for each. If the full profile does not beat the opening sentence on your tasks, the extra length is hard to justify.

A separate paper on how writing style affects LLM safety also examines how persona or framing in a prompt changes model behavior, a reminder that such framing deserves testing rather than assumption.

Frequently asked questions

What is an AGENTS.md file?

An AGENTS.md file is generally described as a plain-text file of standing instructions that coding and general-purpose agents read before working. In the Scientific Agents study, each of the 503 profiles is a profession-specific AGENTS.md file. This is general background rather than a definition from the paper’s abstract, so check an agent’s own documentation for how it handles the file.

Does a detailed system prompt make an AI agent more accurate?

Not consistently, according to the Scientific Agents paper. Its authors report that detailed profession-specific system prompts raised token use and estimated cost per response without a consistent accuracy gain. That is weaker than saying prompts never help, and the abstract gives no figures, so the size of any gain or loss is unknown.

Why does a longer system prompt cost more?

Under token-based pricing, the instructions a model receives count as input tokens, so a longer system prompt costs more on each call that includes it. Agents that resend their instructions repeatedly pay the overhead repeatedly. The Scientific Agents paper describes its cost as an estimate per response, not a billed amount.

Does this apply to models other than Gemini 3.8 Flash?

That is not known yet. The Scientific Agents results come from Gemini 3.8 Flash, run through OpenRouter in the Pi agent harness, and the abstract names no other model. Other models may respond differently to long profiles, so teams should test their own model before changing how they write instructions.