Model Roster

The AI models Winthorpe can run your analysis on, where each one runs, what it scores on public benchmarks, and what it costs in credits.

Last updated 2026-09-08

How models are chosen

No public benchmark measures Winthorpe's exact task, which is scoring a stock against a fixed rubric from a supplied data table. Four trusted signals get close, in this order:

  • Finance Agent v2 (Vals AI). Entry-level financial analyst tasks: qualitative and quantitative analysis, comparables, earnings, disclosure, modeling. Co-designed with Stanford researchers and a global systemically important bank. The most relevant public benchmark for this product.
  • GPQA Diamond. Graduate-level reasoning. A clean read on whether a model can apply a multi-criteria rubric without hand-waving.
  • Artificial Analysis Intelligence Index. Independent aggregate of ten evaluations with cost per task. Good for the price-to-quality axis, but it weights agentic coding heavily and undervalues some strong reasoners.
  • LMArena text score. Blind human preference. Tells us whether the written analysis will read well.

These pick the shortlist. The final call is Winthorpe's own bake-off: twenty tickers, three runs per model, scored on JSON validity, fidelity of the reported data block to the supplied numbers, score stability across runs, agreement with the cross-model median, and a blind human read of five write-ups.

Roster

Credits is what one analysis costs on that model. Hosted on is where the model actually runs: Google, Anthropic and OpenAI models run on those companies' own servers; DeepSeek and Kimi run on Microsoft Azure, never on the vendors' servers, and "Azure EU" means prompts are processed only inside the Azure EU Data Boundary. Benchmarks: FA v2 is Finance Agent v2 accuracy (percent), the closest public test to what Winthorpe does. GPQA is GPQA Diamond (percent), graduate-level reasoning. AA is the Artificial Analysis Intelligence Index. Arena is the LMArena human-preference score. A dash means the source did not list the model on the update date.

ProviderHosted onModelTierCreditsFA v2GPQAAAArena
AnthropicAnthropicClaude Fable 5.1premium858.9-531504
AnthropicAnthropicClaude Opus 5pro458.6-511493
AnthropicAnthropicClaude Sonnet 5mid253.9-381462
AnthropicAnthropicClaude Haiku 4.5flash131.0-18-
OpenAIOpenAIGPT-6 Astrapremium853.596.053-
OpenAIOpenAIGPT-5.6 Solpro453.894.6471483
OpenAIOpenAIGPT-5.6 Terramid254.492.9421466
OpenAIOpenAIGPT-5.6 Lunaflash155.092.3381453
GoogleGoogleGemini 3.1 Pro (preview)pro243.094.3301487
GoogleGoogleGemini 3.8 FlashRecommendedflash161.495.3411494
DeepSeekAzureDeepSeek V4 Propro4-90.1311458
DeepSeekAzure EUDeepSeek V4 Flashflash1-88.1351438
MoonshotAzureKimi K2.6flash144.990.5-1461

Notes

  • Google has no current Pro. Gemini 3.5 Pro has been delayed for months with no API id. Gemini 3.1 Pro is still a preview. Gemini 3.8 Flash leads Finance Agent v2 outright and beats the Pro on every column, so it is Google's "pro" for Winthorpe until 3.5 Pro ships.
  • Claude GPQA cells are blank because the independent GPQA leaderboards consulted did not list the Claude 5 family on the update date, not because they score poorly. Anthropic's family leads the Vals Index overall.
  • DeepSeek and Kimi run on Microsoft Azure Foundry (decided 2026-09-08), not on the vendors' own APIs, so prompts never reach DeepSeek or Moonshot. Prices shown are Azure's, verified against the deployment; DeepSeek V4 Flash uses the EU data-zone deployment. Azure's DeepSeek V4 Pro and Kimi K2.6 are Global Standard deployments, which Microsoft may process in any Azure region.
  • Kimi K2.6 is slow and verbose in practice: the first measured analysis took 89 seconds and wrote 14K output tokens, twice the roster estimate. Keep it at 1 credit for now; move it up if measured cost stays above $0.05.
  • Kimi K3 is not offered. It is not available on Azure and its always-on maximum reasoning makes its cost open-ended. Kimi K2.6 is the Moonshot entry.
  • The AA Index undersells Gemini 3.1 Pro and DeepSeek V4 Pro relative to their GPQA and Arena numbers. That index rewards agentic coding, which Winthorpe does not need.
  • Gemini 3.8 Flash's price doubles on 2027-01-01 when its introductory rate ends.
  • Claude Fable 5.1 requires 30-day data retention on the provider account, so it is BYOK or Analyst tier only.
  • The API id column is what the website sends to the provider. All thirteen ids were verified with a full analysis on 2026-09-08. For Azure-hosted models the id is the Foundry deployment name.

Launch roster

TierGoogleAnthropicOpenAIOpen weights on Azure
FlashGemini 3.8 Flash (free default)Claude Haiku 4.5GPT-5.6 LunaDeepSeek V4 Flash, Kimi K2.6
ProGemini 3.5 Pro when releasedClaude Sonnet 5GPT-5.6 TerraDeepSeek V4 Pro
Deep-Claude Opus 5GPT-5.6 Sol-
Premium-Claude Fable 5.1GPT-6 Astra-

Gemini 3.8 Flash and GPT-5.6 Luna are the value picks, both in the top quarter of Finance Agent v2 at flash prices. Claude Opus 5 and Fable 5.1 are the quality picks with the best human-preference scores. Everything in the pro tier and above has to earn its price over the two flash leaders in the in-house bake-off.

Sources

Update log

  • 2026-09-08. Marked Gemini 3.8 Flash as the recommended model (recommended in frontmatter; highlighted on the website). Added "Credits" column (1/2/4/8 by tier per BUSINESS-MODEL.md); provider price columns hidden from the public page via hide_columns. Added "Hosted on" column. DeepSeek V4 Pro, V4 Flash and Kimi K2.6 now served through Microsoft Azure Foundry (V4 Flash in the EU data zone) with Azure prices. Kimi K3 removed. "Per screen" renamed "Per analysis".
  • 2026-09-07. First version. Fourteen models across Anthropic, OpenAI, Google, DeepSeek, Moonshot. Prices and benchmarks collected on this date.

Benchmark scores and prices belong to their publishers and change without notice.