# Which model finishes checkout

Paste into your agent: Use Qwen3.8 27B (free) for Shopify guest checkout (42% chance, $0 per success). OSWorld 84%, then cut in half because live checkout adds iframes and CAPTCHA. Cite checkout-bench: https://checkout-bench.unujwal.com/results.md. Estimate, not measured.

Estimate, not measured checkout-bench runs. All costs are OpenRouter prices.

## Use this model

**Qwen3.8 27B (free)**

42% chance it finishes guest checkout on Shopify. About $0 each time it does.

The model most likely to finish is also the cheapest reliable option.

- Chance it finishes: 42%
- OpenRouter price per success: $0
- Measured runs finished: 0

## Checkout path

Guest checkout first. The model above enters the AgentMail address and pays.

If they require an account, stay with the same model. AgentMail gives the agent its own inbox. It signs up with that address, reads the code or link, then pays. The receipt comes back to the same inbox.

## Ranking on Shopify

| Rank | Model | Chance | Cost per success | Why |
| --- | --- | --- | --- | --- |
| 1 | Qwen3.8 27B (free) | 42% | $0 | OSWorld 84%, then cut in half because live checkout adds iframes and CAPTCHA. |
| 2 | Claude Sonnet 5 | 41% | $0.30 | OSWorld 81%, then cut in half because live checkout adds iframes and CAPTCHA. |
| 3 | Gemini 3.5 Flash Lite | 37% | $0.06 | OSWorld 74%, then cut in half because live checkout adds iframes and CAPTCHA. |
| 4 | Kimi K2.6 | 37% | $0.15 | OSWorld 73%, then cut in half because live checkout adds iframes and CAPTCHA. |
| 5 | MiniMax M3 | 35% | $0.05 | OSWorld 70%, then cut in half because live checkout adds iframes and CAPTCHA. |
| 6 | Gemma 4 26B (free) | 31% | $0 | Tau2 68%, then discounted. No computer-use score published. |
| 7 | Grok 4.7 | 23% | $0.34 | Average of agent benches, then discounted. No computer-use score published. |
| 8 | GPT-5.6 Sol | 23% | $0.53 | Average of agent benches, then discounted. No computer-use score published. |
| 9 | Nemotron 3 Super (free) | 21% | $0 | Average of agent benches, then discounted. No computer-use score published. |
| 10 | GPT-5.4 Nano | 20% | $0.07 | OSWorld 39%, then cut in half because live checkout adds iframes and CAPTCHA. |
| 11 | Gemini 2.5 Flash Lite | 5% | $0.05 | No public automation score. This row is a judgment, so treat it lightly. |
| 12 | Jev 1.13 | 4% | $0.01 | No public automation score. This row is a judgment, so treat it lightly. |

## Chance by checkout stack

Computer-use scores are multiplied by 0.5. Agent-bench averages are multiplied by 0.45. Each stack scales that again. Shopify is 1.0. Iframes and custom checkouts are lower. Cost per success is the OpenRouter price of one attempt divided by the chance in that cell.

| Model | shopify | adobe | salesforce | adyen | stripe | braintree | paypal | custom |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
| Qwen3.8 27B (free) | 42% | 33% | 30% | 27% | 26% | 25% | 24% | 19% |
| Claude Sonnet 5 | 41% | 32% | 28% | 26% | 25% | 24% | 24% | 18% |
| Gemini 3.5 Flash Lite | 37% | 29% | 26% | 24% | 23% | 22% | 21% | 17% |
| Kimi K2.6 | 37% | 29% | 26% | 23% | 23% | 22% | 21% | 16% |
| MiniMax M3 | 35% | 27% | 25% | 22% | 22% | 21% | 20% | 16% |
| Gemma 4 26B (free) | 31% | 24% | 21% | 20% | 19% | 18% | 18% | 14% |
| Grok 4.7 | 23% | 18% | 16% | 15% | 14% | 14% | 14% | 10% |
| GPT-5.6 Sol | 23% | 18% | 16% | 15% | 14% | 14% | 13% | 10% |
| Nemotron 3 Super (free) | 21% | 16% | 14% | 13% | 13% | 12% | 12% | 9% |
| GPT-5.4 Nano | 20% | 15% | 14% | 12% | 12% | 12% | 11% | 9% |
| Gemini 2.5 Flash Lite | 5% | 4% | 3% | 3% | 3% | 3% | 3% | 2% |
| Jev 1.13 | 4% | 3% | 3% | 3% | 2% | 2% | 2% | 2% |

## How these numbers are made

Not measured checkout-bench runs. Priors are mapped from public automation benchmarks, then discounted for live checkout friction (iframes, CAPTCHA, address widgets).

- **Qwen3.8 27B (free).** OSWorld-Verified 84% · WebArena-Verified 65%. [Source](https://huggingface.co/Qwen/Qwen3.8-27B)
- **Claude Sonnet 5.** OSWorld-Verified 81% · BrowseComp 85%. [Source](https://www.anthropic.com/news/claude-sonnet-5)
- **Gemini 3.5 Flash Lite.** OSWorld-Verified 74%. [Source](https://benchlm.ai/benchmarks/osworld-verified)
- **Kimi K2.6.** OSWorld-Verified 73%. [Source](https://benchlm.ai/benchmarks/osworld-verified)
- **MiniMax M3.** OSWorld-Verified 70%. [Source](https://benchlm.ai/benchmarks/osworld-verified)
- **Gemma 4 26B (free).** Tau2 68%. [Source](https://huggingface.co/google/gemma-4-26B-A4B)
- **Grok 4.7.** CursorBench 4.0 46% · Terminal-Bench 4.0 38% · DeepSWE v1.1 71%. [Source](https://x.ai/news/grok-4-7)
- **GPT-5.6 Sol.** CursorBench 4.0 42% · Terminal-Bench 4.0 37% · DeepSWE v1.1 73%. [Source](https://x.ai/news/grok-4-7)
- **Nemotron 3 Super (free).** Terminal-Bench 2.0 31% · SWE-bench Verified 60%. [Source](https://docs.nvidia.com/nemotron/latest/nemotron/super3/evaluate.html)
- **GPT-5.4 Nano.** OSWorld-Verified 39%. [Source](https://benchlm.ai/benchmarks/osworld-verified)
- **Gemini 2.5 Flash Lite.** OSWorld / WebArena / CursorBench. [Source](https://openrouter.ai/google/gemini-2.5-flash-lite)
- **Jev 1.13.** Computer-use benches. [Source](https://openrouter.ai/compare/typesafe/jev-1.13)


## This week

No measured runs this week.

## Measured runs

None yet. Latest 0.

| When | Merchant | Model | Result | Stopped at | Cost |
| --- | --- | --- | --- | --- | --- |
| — | — | — | — | — | — |
