Model Roster
The AI models Winthorpe can run your analysis on, where each one runs, what it scores on public benchmarks, and what it costs in credits.
Last updated 2026-09-08
How models are chosen
No public benchmark measures Winthorpe's exact task, which is scoring a stock against a fixed rubric from a supplied data table. Four trusted signals get close, in this order:
- Finance Agent v2 (Vals AI). Entry-level financial analyst tasks: qualitative and quantitative analysis, comparables, earnings, disclosure, modeling. Co-designed with Stanford researchers and a global systemically important bank. The most relevant public benchmark for this product.
- GPQA Diamond. Graduate-level reasoning. A clean read on whether a model can apply a multi-criteria rubric without hand-waving.
- Artificial Analysis Intelligence Index. Independent aggregate of ten evaluations with cost per task. Good for the price-to-quality axis, but it weights agentic coding heavily and undervalues some strong reasoners.
- LMArena text score. Blind human preference. Tells us whether the written analysis will read well.
These pick the shortlist. The final call is Winthorpe's own bake-off: twenty tickers, three runs per model, scored on JSON validity, fidelity of the reported data block to the supplied numbers, score stability across runs, agreement with the cross-model median, and a blind human read of five write-ups.
Roster
Credits is what one analysis costs on that model. Hosted on is where the model actually runs: Google, Anthropic and OpenAI models run on those companies' own servers; DeepSeek and Kimi run on Microsoft Azure, never on the vendors' servers, and "Azure EU" means prompts are processed only inside the Azure EU Data Boundary. Benchmarks: FA v2 is Finance Agent v2 accuracy (percent), the closest public test to what Winthorpe does. GPQA is GPQA Diamond (percent), graduate-level reasoning. AA is the Artificial Analysis Intelligence Index. Arena is the LMArena human-preference score. A dash means the source did not list the model on the update date.
| Provider | Hosted on | Model | Tier | Credits | FA v2 | GPQA | AA | Arena |
|---|---|---|---|---|---|---|---|---|
| Anthropic | Anthropic | Claude Fable 5.1 | premium | 8 | 58.9 | - | 53 | 1504 |
| Anthropic | Anthropic | Claude Opus 5 | pro | 4 | 58.6 | - | 51 | 1493 |
| Anthropic | Anthropic | Claude Sonnet 5 | mid | 2 | 53.9 | - | 38 | 1462 |
| Anthropic | Anthropic | Claude Haiku 4.5 | flash | 1 | 31.0 | - | 18 | - |
| OpenAI | OpenAI | GPT-6 Astra | premium | 8 | 53.5 | 96.0 | 53 | - |
| OpenAI | OpenAI | GPT-5.6 Sol | pro | 4 | 53.8 | 94.6 | 47 | 1483 |
| OpenAI | OpenAI | GPT-5.6 Terra | mid | 2 | 54.4 | 92.9 | 42 | 1466 |
| OpenAI | OpenAI | GPT-5.6 Luna | flash | 1 | 55.0 | 92.3 | 38 | 1453 |
| Gemini 3.1 Pro (preview) | pro | 2 | 43.0 | 94.3 | 30 | 1487 | ||
| Gemini 3.8 FlashRecommended | flash | 1 | 61.4 | 95.3 | 41 | 1494 | ||
| DeepSeek | Azure | DeepSeek V4 Pro | pro | 4 | - | 90.1 | 31 | 1458 |
| DeepSeek | Azure EU | DeepSeek V4 Flash | flash | 1 | - | 88.1 | 35 | 1438 |
| Moonshot | Azure | Kimi K2.6 | flash | 1 | 44.9 | 90.5 | - | 1461 |
Notes
- Google has no current Pro. Gemini 3.5 Pro has been delayed for months with no API id. Gemini 3.1 Pro is still a preview. Gemini 3.8 Flash leads Finance Agent v2 outright and beats the Pro on every column, so it is Google's "pro" for Winthorpe until 3.5 Pro ships.
- Claude GPQA cells are blank because the independent GPQA leaderboards consulted did not list the Claude 5 family on the update date, not because they score poorly. Anthropic's family leads the Vals Index overall.
- DeepSeek and Kimi run on Microsoft Azure Foundry (decided 2026-09-08), not on the vendors' own APIs, so prompts never reach DeepSeek or Moonshot. Prices shown are Azure's, verified against the deployment; DeepSeek V4 Flash uses the EU data-zone deployment. Azure's DeepSeek V4 Pro and Kimi K2.6 are Global Standard deployments, which Microsoft may process in any Azure region.
- Kimi K2.6 is slow and verbose in practice: the first measured analysis took 89 seconds and wrote 14K output tokens, twice the roster estimate. Keep it at 1 credit for now; move it up if measured cost stays above $0.05.
- Kimi K3 is not offered. It is not available on Azure and its always-on maximum reasoning makes its cost open-ended. Kimi K2.6 is the Moonshot entry.
- The AA Index undersells Gemini 3.1 Pro and DeepSeek V4 Pro relative to their GPQA and Arena numbers. That index rewards agentic coding, which Winthorpe does not need.
- Gemini 3.8 Flash's price doubles on 2027-01-01 when its introductory rate ends.
- Claude Fable 5.1 requires 30-day data retention on the provider account, so it is BYOK or Analyst tier only.
- The API id column is what the website sends to the provider. All thirteen ids were verified with a full analysis on 2026-09-08. For Azure-hosted models the id is the Foundry deployment name.
Launch roster
| Tier | Anthropic | OpenAI | Open weights on Azure | |
|---|---|---|---|---|
| Flash | Gemini 3.8 Flash (free default) | Claude Haiku 4.5 | GPT-5.6 Luna | DeepSeek V4 Flash, Kimi K2.6 |
| Pro | Gemini 3.5 Pro when released | Claude Sonnet 5 | GPT-5.6 Terra | DeepSeek V4 Pro |
| Deep | - | Claude Opus 5 | GPT-5.6 Sol | - |
| Premium | - | Claude Fable 5.1 | GPT-6 Astra | - |
Gemini 3.8 Flash and GPT-5.6 Luna are the value picks, both in the top quarter of Finance Agent v2 at flash prices. Claude Opus 5 and Fable 5.1 are the quality picks with the best human-preference scores. Everything in the pro tier and above has to earn its price over the two flash leaders in the in-house bake-off.
Sources
Update log
- 2026-09-08. Marked Gemini 3.8 Flash as the recommended model (
recommendedin frontmatter; highlighted on the website). Added "Credits" column (1/2/4/8 by tier per BUSINESS-MODEL.md); provider price columns hidden from the public page viahide_columns. Added "Hosted on" column. DeepSeek V4 Pro, V4 Flash and Kimi K2.6 now served through Microsoft Azure Foundry (V4 Flash in the EU data zone) with Azure prices. Kimi K3 removed. "Per screen" renamed "Per analysis". - 2026-09-07. First version. Fourteen models across Anthropic, OpenAI, Google, DeepSeek, Moonshot. Prices and benchmarks collected on this date.
Benchmark scores and prices belong to their publishers and change without notice.