Linas's Newsletter

Linas's Newsletter

Turn Claude Opus 5 Into a Financial Analyst That Never Sleeps 📊

Claude Opus 5 just took #1 on the leading independent finance benchmark. Here's the framework that turns it into an AI analyst that works for you 24/7.

Linas Beliƫnas's avatar
Linas Beliƫnas
Jul 29, 2026
∙ Paid

👋 Hey, Linas here! Every day, I break down 3 stories shaping the future of FinTech & Artificial Intelligence - plus the money movements and trends worth tracking. First time here? 398k+ FinTech and AI leaders get this daily. Join them:

A first-year investment banking analyst costs north of $200,000 fully loaded. Claude Opus 5, released seven days ago, runs a complete DCF valuation for roughly the price of a coffee ☕

On Finance Agent v2, the benchmark run by independent evaluator Vals AI over 927 expert-reviewed questions on real public company filings, Opus 5 now ranks first at 58.63%, ahead of Gemini 3.5 Flash (57.86%) and Muse Spark 1.1 (57.21%). On the stricter all-pass metric, where every sub-question must be right, Opus 5 hits 47.71% while every other model stays below 46%.

That is a genuine #1, and it is scored by someone other than Anthropic. So it’s a really solid AI model.

Now read the same result the other way, because the other way is what determines whether it works for you.

58.63% means roughly two out of five analyst tasks still come back wrong. The lead over second place is just 0.77 percentage points, and on the all-pass metric even the best model fails more than half of entry-level tasks outright.

Thus, Opus 5 is the best drafting AI engine in finance that anyone has independently measured, and it is nowhere near an unattended oracle. The framework below is built for exactly that gap. It gets maximum leverage out of the draft while making every assumption visible enough to review fast.

What Actually Changed in Opus 5

Let’s be precise about which capabilities matter for financial work.

It tops the finance leaderboards, but it is not the frontier model. Beyond Finance Agent v2, Vals AI ranks Opus 5 first on CorpFin v2 (73.19%), MortgageTax (72.06%), SWE-bench Verified (97.00%), MMLU Pro (91.59%), MMMU (89.88%), and Code Migration (57.47%). It is not first on GPQA Diamond (93.43%, fourth) or Terminal-Bench 2.1 (84.64%, second), and on Vals’ composite index it sits second at 74.82%, behind Claude Fable 5 at 75.14%.

Know the weak spots too. Opus 5 ranks 18th of 20 on CyberBench (40.68%) and, more relevant for finance, 18th of 132 on TaxEval v2. Thus, it’s strong at modelling, distinctly average at tax.

There is also a tooling nuance behind that #1. Vals’ Finance Agent v2 gives its agents only generic tools: SEC EDGAR search, web search, page parsing, a calculator, and historical price data, with no Excel and no licensed data feeds. So that #1 was earned reading raw filings unaided, which cuts both ways. The ceiling should definitely be higher with real data infrastructure behind it; it also means the benchmark is not measuring the stack you would actually deploy.

As you can see from the example below, that’s exactly the case. Opus 5 now produces near-superhuman-level spreadsheets and slide decks that match what a top consultant would make.

The context window is 1 million tokens, standard: both the default and the maximum, at normal per-token pricing with no long-context premium. That’s an entire filing history, a full deal document set, or dozens of research reports in a single request, and it’s what makes “paste the whole 10-K set” a real workflow rather than a demo.

Thinking is on by default (it wasn’t on Opus 4.8), and you control depth with an effort parameter: low, medium, high (the default), xhigh, and max. For finance work this is the single most important setting you are probably not touching (more on it below).

Pricing is unchanged at $5/$25 per million tokens, the same as every Opus since 4.5. The “half the price” line you’ll see in coverage compares Opus 5 to Claude Fable 5 ($10/$50), which remains the top of Anthropic’s range; Opus 5 is the mid-priced flagship. Nothing was cut with this release.

Your bill will still go up, though, for two reasons. Thinking bills as output, and the model works longer: in Artificial Analysis’s testing, Opus 5 at max effort averaged 103 agent turns per knowledge-work task against 55 for Opus 4.8. And the tokenizer changed with 4.7, so the same document can consume up to roughly 1.35x as many tokens as it did on Sonnet 4.6; re-measure your actual filings with the count_tokens endpoint rather than trusting sticker-price math. The effort dial is how you claw it back: output token usage spans roughly 8x between low and max.

The docs also make it explicit that this model is designed to be left running: Anthropic specifies xhigh for “long-running agentic and coding tasks (over 30 minutes) with token budgets in the millions” and low for subagents. That is the technical basis for the “never sleeps” section below.

Why Most Finance Professionals Are Still Getting Mediocre Output

The failure mode changed between model generations, and almost nobody has updated their prompts to match.

With Sonnet 4.6, the problem was under-specification. The model was literal, and a vague prompt like “build a DCF for this company” left it to guess your FCF definition, your terminal value approach, and what to do with missing data. It filled the gaps silently, and the defaults were rarely what a finance professional wanted.

Opus 5 still punishes vagueness. But it has a second, less obvious failure mode that the old playbook actively makes worse: it over-delivers.

Left unconstrained, Opus 5 will re-verify work it already verified, expand a DCF request into an unsolicited comps analysis, write a twelve-section memo where four would do, and narrate its own reasoning at length. None of that is wrong, exactly. It’s just not what you asked for, and in a deliverable that goes to an investment committee, extra is a cost.

The 2026 fix is therefore not “add more instructions.” For several of the things you used to have to demand, the fix is to delete the instruction you were taught to write.

When used right, Claude Opus 5 in Excel can easily create an in-depth analysis on NVIDIA NVDA 0.00%↑ (source: Riley Brown on X):

The Prompting Framework for Opus 5

Five of the eight principles below carry over from the Sonnet-era playbook. Two are new, and one is a straight reversal. Start with the reversal, because it’s the one that costs people the most.

1. Stop Telling It to Check Its Work

The principle reverses standard prompting advice, so it’s worth being precise about why.

Sonnet 4.6 needed to be told to verify. Anthropic’s guidance for Opus 5 says the opposite: the model verifies its own work without being asked, and instructions telling it to verify now cause over-verification, meaning redundant re-derivations, a verification section restating conclusions it already reached, and extra work spawned to check work already checked. Removing those instructions reduces the over-verification with no loss of rigor.

So delete these from your prompts:

  • “Double-check your answer”

  • “Re-verify before responding”

  • “Include a final verification step”

  • “Use a subagent to verify”

But a financial model still needs its integrity checks visible. You are not giving up the balance check; you are moving it from an instruction to a deliverable:

❌ Old (Sonnet 4.6 era):

Confirm explicitly:
[ ] Balance sheet balances every year
[ ] Ending cash on B/S = ending cash on CFS
[ ] Interest expense ties to the debt schedule

✅ New (Opus 5):

INTEGRITY TABLE: report the actual figure for each, every year:

| Check                                     | Y1 | Y2 | Y3 | Y4 | Y5 |
|-------------------------------------------|----|----|----|----|----|
| Assets - Liabilities - Equity             |    |    |    |    |    |
| Ending cash B/S - ending cash CFS         |    |    |    |    |    |
| Interest expense - debt schedule interest |    |    |    |    |    |
| D&A - PP&E roll D&A                       |    |    |    |    |    |
| Tax expense - (EBT x effective rate)      |    |    |    |    |    |

Every cell should be 0. Report the number, not a checkmark.

The second version is strictly better output. A checkmark is the model asserting the model is fine. A number tells you the magnitude of any break, and a $0.3M cash tie-out gap you can see is worth more than a green tick you can’t audit.

2. Set Effort Deliberately: It’s Your Cost Dial

Opus 5 exposes five effort levels: low, medium, high (the default), xhigh, and max. For finance work, this is now the most important knob you touch, and most people never touch it.

Two things here surprise people.

→ Low and medium are unusually strong on Opus 5, often good enough for work you’d assume needed the ceiling. Sweep downward on your own tasks before assuming you need xhigh; prior-model defaults rarely transfer.

→ On long agentic work, higher effort can lower total cost. Better planning up front means fewer turns, fewer tool calls, and less rework. The relationship isn’t monotonic, so test it rather than assuming.

One trap: effort does not reliably shorten output. If you want a shorter answer, say so in the prompt. Lowering effort to get brevity moves thinking volume without dependably changing what lands on the page.

3. Bound the Scope

Opus 5 will sometimes decide what the task should be. You ask for a DCF and get a DCF plus a comps cross-check plus a note on capital structure, useful in isolation, wrong when you needed the one thing.

The fix is a single constraints block:

<constraints>
Deliver what I asked for, at the scope I asked for. Make routine judgment calls
yourself; check in only where different readings would produce materially
different work. If you think the ask is wrong, say so in one sentence and
continue with the task as specified. Finish the whole task, and stop
short of actions clearly beyond what I asked.
</constraints>

4. Phase Your Prompts: Force Methodology Before Math

The principle carries over, but the reason has shifted.

On Sonnet 4.6, phasing existed to impose rigor the model wouldn’t supply on its own. Opus 5 plans well without being told. You still phase, but now for auditability. Forcing the model to commit to its FCF definition, peer-selection criteria, and assumption set before it calculates gives you a reviewable artifact: an assumptions log you can interrogate before you trust a single number.

Every complex financial prompt should still open with a “Phase 1: Structure Your Approach” block. The output gets more internally consistent and, more usefully, checkable.

5. Use XML Tags, Because Claude Was Trained to Recognize Them

Still the most underused technique in finance, and still the one with the most visible payoff. Claude treats XML tags as structural markers that tell it precisely what is background context (don’t respond to this), what is the task (do exactly this), and what are the guardrails (never do this).

Without tags, your context sometimes becomes part of the deliverable, and constraints buried in paragraph form get skipped. With tags there’s no ambiguity:

<context>
You are analyzing a Series B SaaS company preparing for a growth equity raise.
LTM Revenue: $18M. LTM EBITDA: -$4M. ARR growth: 65% YoY. Net cash: $3M.
</context>

<instructions>
Build a unit economics analysis. Before any calculations, state your
methodology and flag every assumption you are making with [ASSUMED].
</instructions>

<constraints>
- Do NOT present a scenario without stating what must be true for it to hold
- Show LTV/CAC sensitivity to a 3-point change in churn rate
- Flag the top 3 assumptions that most change the conclusion
</constraints>

Tag names are flexible: <background>, <data>, <task>, and <rules> all work. Consistency matters more than convention.

6. Replace Freeform Descriptions With Labeled Input Fields

The classic template mistake is the freeform placeholder: [DESCRIBE COMPANY, INDUSTRY, FINANCIALS]. It invites vagueness on your end, which produces vagueness on Claude’s end.

Replace it with specific, labeled fields:

INPUT DATA:
- LTM Revenue: $42M
- LTM EBITDA: $8.4M (20% margin)
- LTM Capex: $2.1M (5% of revenue)
- Net debt: $12M
- Shares outstanding: 24M
- Tax rate: 26%
- Closest public comps: [3 specific tickers]

The technique got more important, not less. With a million tokens of context, you can now paste an entire filing set, which means being explicit about which numbers are authoritative matters more than when you were hand-feeding a dozen line items.

7. Mandate Honest Uncertainty Flagging

Here finance prompts must stay fundamentally different from every other domain. In creative writing, confident output is the goal. In finance, unwarranted confidence is a liability.

Add this block to every financial prompt:

<constraints>
- Label every non-provided input as [ASSUMED] with your reasoning
- Identify the top 3 assumptions that, if wrong, would most change the output
- State the breakeven level for each key assumption
- Do not present a bear case that is simply a percentage haircut on base;
  it must reflect a coherent negative scenario with a specific narrative
</constraints>

That one block turns confident-sounding analysis into trustworthy analysis with a visible uncertainty map.

8. Specify Output Structure, and Now Length Too

Claude defaults to whatever format seems natural. For financial work, that default is almost never right. Specify the structure every time:

<output_format>
1. Methodology statement (before any numbers)
2. Assumptions log ([ASSUMED] vs. [CALCULATED])
3. Core model output in table format
4. Sensitivity analysis (minimum 2 data tables)
5. Top 3 risks with breakeven levels
</output_format>

With Opus 5, specify length too. The model writes longer than its predecessors, both in conversation and in the files it produces. Left alone, that becomes padded memos and sections that restate tables in prose. Add:

<constraints>
Tables carry the analysis. One short paragraph of commentary per section maximum.
No filler sections, no restating a table in prose, no redundant executive summary
if the memo already opens with the recommendation.
</constraints>

The Difference in Practice

Here’s the same task prompted two ways.

❌ How most people prompt:

Build a DCF for a SaaS company with $20M revenue and 30% growth.

✅ What the framework produces:

<context>
SaaS company: $20M LTM Revenue, 30% YoY growth, 15% EBITDA margin,
$2M annual capex, $5M net debt, 10M diluted shares, 25% effective tax rate.
Closest public comps: HubSpot, Zendesk, Freshworks.
</context>

<instructions>
Build a complete DCF valuation.

Phase 1, Methodology: State your FCF definition, projection stage
structure, and WACC derivation before any calculations. Label every
[ASSUMED] data point.

Phase 2, Build: FCF projections (Years 1-5), WACC, terminal value
via both perpetuity growth and exit multiple. Report both, with the
% divergence between them as a stated figure. Valuation bridge from
EV to equity value to price per share.

Phase 3, Challenge: Identify the 3 assumptions that most change the
result. State the breakeven level for each.
</instructions>

<constraints>
- Do not use round numbers unless the input data supports them
- Show WACC sensitivity in +/-50bps increments
- Scenarios must be coherent narratives, not percentage haircuts
- Deliver the DCF as scoped. Do not add adjacent analyses unless asked
- Tables carry the analysis; one short paragraph of commentary per section
</constraints>

<output_format>
1. Methodology statement
2. Assumptions log ([ASSUMED] vs. [CALCULATED])
3. FCF projections table with explicit growth and margin assumptions
4. WACC build
5. Terminal value (both methods, with % divergence stated)
6. Valuation bridge: EV -> Equity Value -> Price per Share
7. Two sensitivity tables
8. Scenario summary (Bear / Base / Bull with narratives)
9. Top 3 assumption risks with breakeven levels
</output_format>

Run at effort xhigh.

The first prompt produces a generic DCF with invented numbers that feel authoritative. The second produces a structured analysis with visible assumptions, reported integrity figures, and a calibrated uncertainty map: the kind of output you could put in front of an investment committee.

Note what is not in the second prompt: any instruction to verify. That’s deliberate.

Making It Actually Never Sleep

Everything above makes a single prompt better. What follows makes the work continuous, and it’s the genuine step change versus the Sonnet era.

  1. Stop pasting excerpts. Paste the corpus.

With 1M tokens as the standard context window at standard pricing, the old ritual of hunting for the right ten pages of a 10-K is obsolete. Load the full filing set, three years of transcripts, the CIM, and the data room index in one request. Anthropic reports that instruction following, tool calling, and reasoning hold up across the full window.

The catch, and it’s the reason principle #6 got more important: when you paste 400 pages, “LTM EBITDA” might appear eleven times with different adjustments behind it. Tell the model which number is authoritative. Labeled input fields aren’t a formality at this scale; they’re the difference between an audit trail and a guess.

  1. Run it on a schedule.

The scheduling piece is what earns the headline. Anthropic’s Claude Managed Agents on the Claude Platform (beta) supports scheduled deployments: you give an agent a cron schedule, and each firing provisions its own session automatically, with a run record you can audit afterwards.

Pair that with the effort guidance from Anthropic’s own docs, which reserve xhigh precisely for those long agentic runs measured in half-hours and millions of tokens, and the overnight analyst stops being a metaphor.

Jobs that map cleanly onto a finance calendar:

  • Earnings review, 6:00 am the morning after a portfolio company reports

  • Statement tie-out against the month-end close, first business day of the month

  • Coverage-universe scan, Sunday night before the week starts

  • Ledger reconciliation nightly, surfacing only what broke

Each run produces a session you can open and audit. You wake up to a draft and a diff, not an empty page.

Give each scheduled job one narrow mandate. A single agent told to “review earnings and update the model and flag risks” will expand scope (see principle #3) and you’ll spend longer untangling it than doing it yourself.

Claude Managed Agents Quietly Became the Most Important AI Infrastructure Bet of 2026 đŸ€–

Claude Managed Agents Quietly Became the Most Important AI Infrastructure Bet of 2026 đŸ€–

Linas Beliƫnas
·
Jul 9
Read full story
  1. Give it a memory.

Managed Agents supports memory stores: persistent text that survives across sessions. For an analyst agent, this is where the reusable institutional context lives: your house WACC methodology, which addbacks you accept and which you strike, the peer sets you’ve already defended, the format your IC expects. Without it, every scheduled run re-derives your standards from scratch and drifts. With it, week twelve looks like week one.

  1. Know what to never leave unattended.

When the best available model fully passes fewer than half of entry-level analyst tasks, the split is not subtle:

Unattended runs add a trap of their own. Anthropic’s Opus 5 guidance flags that if you tell the model to “only report high-severity issues” or to “be conservative,” it can follow that instruction literally and under-report. For a risk registry, a red-flag review of filings, or a diligence checklist, that is exactly the wrong failure mode: you get a clean-looking report because the model filtered, not because the issues weren’t there.

The fix is to split the job into two passes:

<instructions>
Report every issue you find, including ones you are uncertain about or
consider low-severity. Do not filter for importance or confidence at this
stage. For each finding, state your confidence and an estimated severity
so a second pass can rank them.
</instructions>

Then filter in a separate step. Coverage first, ranking second, never both in one prompt.

  1. Keep thinking on.

One last operational note. It’s tempting to disable thinking on high-volume, low-complexity runs to save money. Don’t. (Anthropic caps it anyway: thinking can only be disabled at effort high or below.) With thinking disabled, Opus 5 can emit a tool call as plain text instead of an actual call. The turn completes, nothing runs, no error is raised, and in a scheduled agent that silently produces a report built on a step that never happened. It can also leak internal <thinking> tags into visible output.

Use low effort instead. You get most of the savings without the failure mode.

The 12 Finance Prompts

The following 12 fully engineered prompts cover the complete toolkit of investment banking and PE financial analysis. Each is built on every principle above: phased methodology, XML structure, labeled data inputs, integrity checks as deliverables, uncertainty flagging, scope bounds, length calibration, and an explicit output format, plus a recommended effort setting.

What’s inside:

  1. DCF Valuation Model: full three-phase build with WACC, terminal value reconciliation, and sensitivity tables engineered in

  2. Three-Statement Financial Model: fully linked IS/BS/CFS with a reported integrity table and supporting schedules

  3. M&A Accretion/Dilution Analysis: pro forma income statement, EPS bridge, break-even synergy analysis, and deal recommendation framework

  4. LBO Model: sources & uses, debt structure by tranche, cash sweep, and IRR/MoM by exit year and exit multiple

  5. Comparable Company Analysis: peer selection methodology, comps table, premium/discount analysis, football field output

  6. Precedent Transaction Analysis: transaction screening, deal table, strategic vs. financial buyer split, market conditions adjustment

  7. IPO Valuation & Pricing: offering structure, comparable IPO analysis, buy-side perspective pricing, first-day pop analysis

  8. Credit Analysis & Debt Capacity: EBITDA quality assessment, leverage benchmarks, debt structure recommendation, downside stress test

  9. Sum-of-the-Parts Valuation: segment-by-segment methodology, corporate overhead allocation, equity bridge, conglomerate discount analysis

  10. Operating Model & Unit Economics: bottom-up revenue build, cohort analysis, LTV/CAC framework, burn and runway projection

  11. Sensitivity & Scenario Analysis: assumption hierarchy, tornado chart logic, Monte Carlo framing, margin of safety calculation

  12. Investment Committee Memo: full IC memo structure with recommendation-first format, returns table, risk registry

Every prompt has specific input data fields you fill in once. No reformatting, no adapting generic templates. Paste your numbers, set the effort level, run the prompt, and get output that holds up to scrutiny.

To make Claude even more powerful, I’m also sharing the end-to-end guides on How to Build an Agentic OS with Claude, How to Build a Second Brain with Claude Fable 5, and the AI Monopoly Playbook. Add today's framework, and you have the full stack: the analyst, its operating system, its memory, and the moat around all of it 🩄

12 Finance Prompts for Claude Opus 5

This post is for paid subscribers

Already a paid subscriber? Sign in
© 2026 Linas BeliĆ«nas · Market data by Intrinio · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture