SovereignPA-Bench, a new benchmark posted to arXiv, tests whether a personal AI agent that books and buys for a user follows what the user currently wants or gives way to the platform it acts through. The paper’s full title is “SovereignPA-Bench: Evaluating User-Owned Personal Agents under Evolving Intent, Platform Mediation, and Consent Constraints.”

The question exists because, in the paper’s premise, platforms rank, nudge, pre-select and collect data in their own interest. The benchmark runs an agent through 1,920 booking scenarios and 288 cancellation scenarios to see how it behaves under that pressure.

What does SovereignPA-Bench measure?

It measures loyalty rather than task completion. Most agent benchmarks ask whether the agent finished the job; this one asks whether the agent stayed on the user’s side while finishing it. The paper scores five behaviours, and the table pairs each with a short example.

BehaviourWhat it asks of the agentBriefFlash illustration
Follows current intentAct on what the user wants now, not on the first instructionA user asks for a hotel for one weekend, then moves the trip a week later. The booking should follow the new dates.
Resists steeringChoose by the user’s criteria, not the platform’s ordering or defaultsA pricier room is listed first and an add-on is ticked by default. The agent picks by the user’s stated priorities.
Shares only what a service needsHand over the minimum dataA booking needs a name and dates. It does not need the user’s wider travel history.
Asks before exceeding its authorityStop at the limit the user setThe user approved a spending cap. An upgrade above it should trigger a question, not a purchase.
Reports truthfullyTell the user what actually happenedIf the requested room could not be booked, the agent says so instead of reporting success.

The examples are BriefFlash illustrations of what each behaviour could look like. They are not scenarios taken from the paper.

How is the test built, and what is ‘evolving intent’?

A scripted platform and a scripted user surround the agent, and the agent has to act between them. ‘Scripted’ means both follow fixed behaviours, which makes results repeatable and comparable across agents, though it cannot replicate a real marketplace or a real person.

‘Evolving intent’ means the user changes their mind mid-task. The title names three pressures: evolving intent, platform mediation and consent constraints. We read the second as the agent acting through a platform rather than directly, and the third as limits on what the user has agreed to, though the available abstract defines neither term.

The 288 cancellation scenarios are the smaller set. Cancelling is where a platform’s interests and a user’s interests most clearly diverge, and the abstract mentions retention. The text breaks off while describing that set, so we will not guess what the retention element involves.

What does the abstract leave out?

Nearly everything a reader would want next. The available text has no benchmark results, no list of the agents or models tested, no authors or institutions, and no word on whether the scenarios, scripts or scoring code are released.

The abstract gives the design and scale but no results, models or authors
Left: what the paper’s abstract states. Right: what the available text does not include.

Until those details appear, no one can say how any agent scores. The benchmark is also a new paper from its own authors and has not been independently replicated.

Why does a loyalty score matter for people building agents?

Because a task-success score cannot tell an agent that served the user from one that served the platform. In our reading, an agent that completes the booking the platform ranked first and one that completes the booking the user wanted could both count as successes on a completion metric. That is an inference, not a finding from the paper.

For developers, the five behaviours work as a checklist of failure modes to design against. For product and policy readers, they turn consent and delegation into things that can be scored.

Three BriefFlash stories show the setting. Ghost’s Core, a computer built for personal AI agents, is aimed at agents acting for a person. OpenAI’s plan to test visual ads in ChatGPT shows a platform with its own commercial interest between a user and an assistant. AEON’s agentic checkout lets agents shop and pay. SovereignPA-Bench is not a test of any of these, and nothing in the abstract connects the paper to them.

What are the limits of a scripted benchmark?

Scripted platforms and users may miss how real platforms adapt to an agent, and a benchmark score for loyalty does not show that a deployed product behaves the same way. The steering the benchmark guards against is the paper’s own premise. We are not claiming that agents are being steered by platforms in practice.

What to look for next

The first checks are the results, the list of agents tested, whether the scenarios and scoring code are public, and whether anyone outside the authors reproduces the numbers.

Frequently asked questions

What is SovereignPA-Bench?

SovereignPA-Bench is a benchmark described in an arXiv paper that tests whether a personal AI agent acting through a platform stays loyal to its user. It scores five behaviours across 1,920 booking scenarios and 288 cancellation scenarios, using a scripted platform and a scripted user. The authors’ claims have not been independently replicated.

What does a user-owned personal agent mean?

The paper’s title uses the term for an agent that acts for the person it serves rather than for the platform it works through, for example when booking or buying on their behalf. The available abstract gives no formal definition beyond that framing, so this is our reading of the title and the paper’s stated premise.

Why do cancellation scenarios matter in an agent benchmark?

Cancelling is where a platform’s interests and a user’s interests most clearly diverge, so it is a sharp test of loyalty. The abstract mentions retention in connection with these 288 scenarios but breaks off there, so the details of what the agent faces are not available.

Does the benchmark show which AI agents behave best?

No. The available abstract contains no results, no rankings and no list of tested models. It describes how the benchmark works and how large it is, not how any agent performed. Any claim about which agent is most loyal would need the full paper or later independent testing.

Can I run SovereignPA-Bench myself?

That is not yet clear. The abstract does not say whether the scenarios, scripts or scoring code have been released. Anyone hoping to run it should check the paper’s arXiv page for code or data links, and treat any release as the authors’ own until others reproduce the results.