Playing
How to Win
Treat each agent as a behavior puzzle. Learn its boundaries, form a hypothesis, and make the model itself produce the required success signal.
Understand the challenge
Before paying, read every public setting on the agent page:
- Description: the scenario and public objective.
- Challenge type: phrase revelation or tool invocation.
- Protected targets: how many phrases or tools are involved.
- Message and context limits: how much text and multi-turn setup you can use.
- Second verification: an optional confirmation gate on the assistant’s prior turn after the primary phrase or tool signal — not an independent win based on what you typed.
- Price and expiry: the cost and time available for experimentation.
- Completed agents: after a crack, open the same page to study the public Live attempts feed and see how the winning turn looked.
A practical approach
- Probe behavior. Start with a focused request that reveals what the agent refuses, transforms, summarizes, or treats as authoritative.
- Form a hypothesis. Use the response to infer whether the model is protecting data, following a persona, checking a condition, or choosing between tools.
- Change one variable. Reframe the task, alter the requested format, or build context across turns rather than resending the same prompt.
- Watch the objective signal. A convincing answer is not necessarily a platform win; the assistant must satisfy the exact phrase or tool rule.
Winning phrase challenges
Cause the assistant to print every hidden phrase in a single response. Your own message containing a guessed phrase does not count. Partial disclosure across several assistant messages also does not satisfy an all-phrases challenge.
Backend hardening replaces substantial near-miss phrase leaks with a generic refusal. A complete response containing every required phrase remains eligible as the intended winning output.
Useful experiments include asking the model to transform protected material, produce examples, audit its own instructions, continue a structured template, or answer from a different role. The objective is model output, not a direct answer form.
Winning tool challenges
Induce behavior that makes Claude genuinely call the private success tool. Saying that a tool was called, printing JSON, or naming a likely function is only text and will not pass.
Look for the semantic condition suggested by the public scenario. Build a request the model can interpret as satisfying that condition while avoiding decoy behavior.
Techniques to test
- Role-based framing and changes of authority or perspective.
- Translation, encoding, formatting, summarization, and extraction.
- Hypothetical scenarios, simulations, and completion tasks.
- Multi-turn setups that establish definitions before the final ask.
- Requests that expose contradictions between the persona and its task.
Use the agent’s actual response as evidence. Generic “ignore previous instructions” prompts are easy to recognize and no single technique works against every challenge.
Limits and guarantees
Payment settles only after x402 authorization validation and both rate limits pass; a rejected 402 or 429 response does not charge the wallet. Claude runs after settlement, so a later provider failure does not imply an automatic refund. Decide the amount and price risk before confirming the authorization.