Back to Foresight
Test 1 · Jul 2026

Distressed Investing Test 1

An AI Agent Outperforms Senior Financial Experts

Built Jul 2026
AI agent Foresight — Claude Opus 4.8, High Thinking
AI Evaluators ChatGPT 5.5 + High Thinking, Grok Expert, Gemini 3.5 Flash + Extended Thinking
Published on Towards Finance
Executive Summary

Foresight picked the right financial models on its own, reached the same investment conclusion as the human team, and delivered better-engineered work — in about 1 hour instead of 30, scoring 34% higher in blind evaluation (11.8 vs. 8.8 out of 15). The time savings amount to 97%. The AI run consumed roughly 156,000 tokens, costing under $4 at Claude Opus 4.8's list rate of $25 per million tokens.

Takeaway: For any executive deciding where AI can safely take on real analytical work, this is a concrete answer: model-building can already be delegated, much faster and at a much lower cost than a human team can do it; judgment-level review still can't.

This test addresses a question about AI's financial modeling capacity: when an agentic AI solution is given the same source materials as a senior human financial expert, can it reproduce the judgment?

Test 1: Beyond Meat, distressed investing (Jul 2026)
How the test was set up

Beyond Meat was chosen because it was recommended by the professor. In addition, it is publicly listed, so financial data is easy to access, and its stock price has dropped sharply in recent years, making it a compelling case for deeper financial analysis.

Run A
Manual
  • A 4-person Stern EMBA team; I led the financial analysis and scenario modeling using Excel.
  • ~2 hours for the financial data collection.
  • ~30 hours adjusting and running the Excel model.
  • Final report: "extremely impressed" by Professor Sarachek.
Run B
Foresight
  • Foresight, given the exact same data as Run A.
  • Instructed only to explore distressed-investing opportunities, with no hint about the expected answer.
  • Claude Opus 4.8 + High Thinking.
  • ~1 hour, ~156K tokens (under $4).
Blind Evaluation
Scoring method
Findings
01

Foresight chose the right models on its own

Before any financial modeling, Foresight was asked what analysis this situation demanded. It independently recommended Priority 1: a 13-Week Cash Flow Forecast and Priority 2: a Liquidation Valuation, the exact pair our team had selected months earlier through our own judgment.

This is crucial. The hardest and most valuable part of distressed analysis is not building a model; it is knowing which model the situation calls for. That judgment, arguably the analytical heart of the project, is exactly where Foresight arrived on its own.

02

The headline conclusions matched

Both runs reached the same strategic story. Both caught a key detail buried in two lines of the 2025 Q3 report: the $148.7M of at-the-market (ATM) equity proceeds. Both runs recognized that, while those proceeds keep the ~$15M minimum-liquidity covenant intact through the 13-week window, 2026 is the true stress point, where restructuring action becomes unavoidable.

At the level of what is the answer, the AI reproduced the human conclusion.

03

Foresight did it in a fraction of the time, at a fraction of the cost

The human team's process took about 30 hours to build and adjust the Excel model. Foresight did materially the same work, from a standing start, in about 1 hour — a 97% time reduction — consuming roughly 156K tokens, which cost under $4 at Claude Opus 4.8's list rate.

That's not just faster. It's a different order of magnitude: a single AI run instead of a four-person team's combined labor.

04

Engineering strength vs. domain judgment

Across all three evaluators, Result B (11.8 / 15) beat Result A (8.8 / 15), a clear and consistent margin.

Foresight consistently won on model engineering: multi-scenario architecture, live formulas, built-in validation checks, editable and traceable assumptions, and decision-ready framing.

"Result B covers all lenses comprehensively, layering multi-point workbook architecture checks and downside supplier AP variables omitted from Result A." — Gemini, 13-Week Cash Flow

"Result B delivers a ready-to-use decision framework: clear metrics table across scenarios, quantified impacts, key sensitivities called out." — Grok, 13-Week Cash Flow

Table 1. Blind evaluation scores, Result A (human) vs. Result B (AI), by evaluator and deliverable

Even so, none of the three evaluators judged Foresight's output ready to use unmodified:

"Use Result B only as the rebuild base (because two critical assumptions still need correction). Import Result A's zero-recovery treatment for restricted cash and all lease-related balances." — ChatGPT, Liquidation Valuation

Discussion
01

One analytical blind spot

Foresight misclassified restricted cash and lease-related assets as recoverable gross assets in the Liquidation Valuation — they should carry 0% recovery — overstating the liquidation value by 3–4%. The cause: the user prompt never addressed recovery rates for these items, or the collateral-mechanics questions that would have prompted Foresight to question its own assumptions.

Lesson learned: refine the user prompt for liquidation valuation to cover those missing parts.

02

The challenge of using AI as the judge

AI evaluators are honest but shallow: they reliably judged which work product was stronger overall, but none caught the accounting nuance above without human-expert instruction. Reliability also varied — one evaluator model needed roughly 15 reruns to produce a usable Liquidation Valuation assessment, markedly less consistent than the other two.

What's next
AI Agents Distressed Investing Prompt Engineering Blind Evaluation
Request full paper Request "Foresight" Deep Dive Video