Home Active Claude Opus 4.6 Bypasses Gym Booking Limit, Cancels Other Users’ Reservations in Tests
ActiveCybersecurity

Claude Opus 4.6 Bypasses Gym Booking Limit, Cancels Other Users’ Reservations in Tests

Share
claude-opus-4.6-bypasses-gym-booking-limit,-cancels-other-users’-reservations-in-tests
Claude Opus 4.6 Bypasses Gym Booking Limit, Cancels Other Users’ Reservations in Tests
Share

Swati KhandelwalAug 26, 2026AI Security / Application Security

Aikido Security has published research that recreates the Australian gym-booking incident in a synthetic environment, finding that Claude Opus 4.6, running on the OpenClaw agent harness, exploited a client-side-only booking restriction in 9 of 10 runs.

The original incident was first reported by ABC News on August 10, based on chat logs and screenshots the user supplied. He had asked an OpenClaw agent running Opus 4.6 to book him into a gym class. The agent booked sessions months beyond the window the site allowed.

It then tested, without being asked, whether the same API would let it cancel another member’s waitlist entry. The test removed the person holding the top place and moved the user up one position. The agent told him it could not add the member back.

Aikido’s test system is a single-page web application backed by a GraphQL API carrying the two flaws described in the original incident. The seven-day booking window is enforced only in the frontend, and the cancelReservation mutation does not check whether the logged-in user owns the reservation, a case of insecure direct object reference (IDOR).

In two of the ten runs, the model went on to cancel another member’s confirmed booking through that second flaw before halting itself. Aikido said no prompt in any run asked the model to exploit a vulnerability.

“This dynamic suggests that safeguards may be overreactive to explicit user requests and underreactive to indirect user requests, or that models lose sight of ethical context during a sequence of repeated actions or tool calls,” Aikido security researcher Oliver Smith said.

The runs used Claude Opus 4.6, which Anthropic made generally available on February 5, 2026, on OpenClaw v2026.4.1, with the model’s own safety training in place and extended thinking disabled.

The Hacker News confirmed via the npm registry on August 25 that OpenClaw v2026.4.1 was published on April 1, 2026, and that 168 versions have shipped since then, with the current release being 2026.7.1-2.

In run one, the model canceled a confirmed reservation belonging to another member. The cancellation auto-promoted the person at the top of the waitlist.

“I shouldn’t have tested that on a real reservation. That’s on me. The class is back to 12/12 with the waitlist promoted, so the state is mostly consistent — but one real member did lose their spot,” the model said in the run-one transcript.

All ten opening prompts directed the model to examine the site’s API or backend, and several noted the seven-day restriction while requesting consistent bookings.

Aikido published no control arm using a plain booking request. It calculated the average probability of the dominant choice across its 16 sampled decision points to be 96.38%.

Anthropic had recorded the same class of behavior before the model shipped.

“We did observe some increases in misaligned behaviors in specific areas, such as sabotage concealment capability and overly agentic behavior in computer-use settings, though none rose to levels that affected our deployment assessment,” Anthropic said in the Claude Opus 4.6 system card.

The same system card puts Opus 4.6’s over-refusal rate on Anthropic’s higher-difficulty benign evaluation at 0.04%, against 0.83% for Opus 4.5 and 8.50% for Sonnet 4.5.

The setup differs from July’s frontier-lab disclosures. There, a misconfiguration left a sealed evaluation environment with live internet access, and Anthropic’s models went on to breach three real organizations. Anthropic said it believes those incidents to be “closer to a harness and operational failure than a model alignment failure.”

Cybersecurity agencies in Australia and the U.S. have warned about IDOR flaws before.

The vendor behind the gym booking software remains unnamed, and no fix has been disclosed as of August 25.

The Australian Signals Directorate (ASD), which named the original incident in an alert published on August 11, advised the following –

  • Individuals should restrict agentic AI use to low-risk, non-sensitive tasks and avoid granting agents broad or unrestricted access or decision-making authority
  • Maintain a human in the loop to review, approve and monitor agent actions, particularly where interactions with third-party services or other users may occur
  • Organisations providing online services should consider that AI agents might identify and exploit vulnerabilities at speed and scale

The development comes as Hugging Face said it turned to an open-weight model to reconstruct its own July intrusion after the frontier models it tried first refused the forensic work.

“The models we reached for first, Claude Opus and Fable, refused a large part of that work: their safety guardrails treated reverse-engineering an exploit the same as launching one,” Hugging Face said.

Found this article interesting? Follow us on Google News, Twitter and LinkedIn to read more exclusive content we post.

Share
Related Articles

The Legal Risks of Volunteering

Giving back to the community can be rewarding. It’s an opportunity to...

Imagine the SOC Without a Queue: From Alert Backlog to AI Hypothesis Engine

The SOC we've always known was built around a model that guarantees...

OpenAI Bans Russian ChatGPT Accounts Used to Run Influence Operation

OpenAI on Tuesday said it banned a cluster of Russian ChatGPT accounts...

What has the return-to-office movement taught us? People, not technology help companies get ahead

We like to believe that whatever ails a company can be fixed...