Foresight picked the right financial models on its own, reached the same investment conclusion as the human team, and delivered better-engineered work — in about 1 hour instead of 30, scoring 34% higher in blind evaluation (11.8 vs. 8.8 out of 15). The time savings amount to 97%. The AI run consumed roughly 156,000 tokens, costing under $4 at Claude Opus 4.8's list rate of $25 per million tokens.
- Foresight autonomously selected the same two financial models that the senior human expert team had chosen through professional judgment.
- The time savings amount to 97%. The AI run consumed roughly 156,000 tokens, costing under $4 at Claude Opus 4.8's list rate of $25 per million tokens.
- The scoring itself was conducted by three independent AI evaluators, working blind and with no knowledge of which work was human and which was AI.
Takeaway: For any executive deciding where AI can safely take on real analytical work, this is a concrete answer: model-building can already be delegated, much faster and at a much lower cost than a human team can do it; judgment-level review still can't.
This test addresses a question about AI's financial modeling capacity: when an agentic AI solution is given the same source materials as a senior human financial expert, can it reproduce the judgment?
- A second-year EMBA "Turnaround, Restructuring, and Distressed Investing" course taught by Professor Joseph Sarachek.
- Assignment: present a distressed investment opportunity and a proposed strategy.
- We identified cash burn as the company’s most critical risk and built two models: a 13-Week Cash Flow Forecast to test near-term liquidity survival, and a Liquidation Valuation to estimate creditor recovery in a downside.
Beyond Meat was chosen because it was recommended by the professor. In addition, it is publicly listed, so financial data is easy to access, and its stock price has dropped sharply in recent years, making it a compelling case for deeper financial analysis.
- A 4-person Stern EMBA team; I led the financial analysis and scenario modeling using Excel.
- ~2 hours for the financial data collection.
- ~30 hours adjusting and running the Excel model.
- Final report: "extremely impressed" by Professor Sarachek.
- Foresight, given the exact same data as Run A.
- Instructed only to explore distressed-investing opportunities, with no hint about the expected answer.
- Claude Opus 4.8 + High Thinking.
- ~1 hour, ~156K tokens (under $4).
- To measure quality objectively, Run A and Run B were compared under blind, source-neutral labels.
- Each was scored by three independent AI evaluators — ChatGPT 5.5 + High Thinking, Grok Expert, and Gemini 3.5 Flash + Extended Thinking.
- Each evaluator scored both work products on Accuracy, Completeness, and Actionability.
Foresight chose the right models on its own
Before any financial modeling, Foresight was asked what analysis this situation demanded. It independently recommended Priority 1: a 13-Week Cash Flow Forecast and Priority 2: a Liquidation Valuation, the exact pair our team had selected months earlier through our own judgment.
This is crucial. The hardest and most valuable part of distressed analysis is not building a model; it is knowing which model the situation calls for. That judgment, arguably the analytical heart of the project, is exactly where Foresight arrived on its own.
The headline conclusions matched
Both runs reached the same strategic story. Both caught a key detail buried in two lines of the 2025 Q3 report: the $148.7M of at-the-market (ATM) equity proceeds. Both runs recognized that, while those proceeds keep the ~$15M minimum-liquidity covenant intact through the 13-week window, 2026 is the true stress point, where restructuring action becomes unavoidable.
At the level of what is the answer, the AI reproduced the human conclusion.
Foresight did it in a fraction of the time, at a fraction of the cost
The human team's process took about 30 hours to build and adjust the Excel model. Foresight did materially the same work, from a standing start, in about 1 hour — a 97% time reduction — consuming roughly 156K tokens, which cost under $4 at Claude Opus 4.8's list rate.
That's not just faster. It's a different order of magnitude: a single AI run instead of a four-person team's combined labor.
Engineering strength vs. domain judgment
Across all three evaluators, Result B (11.8 / 15) beat Result A (8.8 / 15), a clear and consistent margin.
Foresight consistently won on model engineering: multi-scenario architecture, live formulas, built-in validation checks, editable and traceable assumptions, and decision-ready framing.
"Result B covers all lenses comprehensively, layering multi-point workbook architecture checks and downside supplier AP variables omitted from Result A." — Gemini, 13-Week Cash Flow
"Result B delivers a ready-to-use decision framework: clear metrics table across scenarios, quantified impacts, key sensitivities called out." — Grok, 13-Week Cash Flow
Even so, none of the three evaluators judged Foresight's output ready to use unmodified:
"Use Result B only as the rebuild base (because two critical assumptions still need correction). Import Result A's zero-recovery treatment for restricted cash and all lease-related balances." — ChatGPT, Liquidation Valuation
One analytical blind spot
Foresight misclassified restricted cash and lease-related assets as recoverable gross assets in the Liquidation Valuation — they should carry 0% recovery — overstating the liquidation value by 3–4%. The cause: the user prompt never addressed recovery rates for these items, or the collateral-mechanics questions that would have prompted Foresight to question its own assumptions.
Lesson learned: refine the user prompt for liquidation valuation to cover those missing parts.
The challenge of using AI as the judge
AI evaluators are honest but shallow: they reliably judged which work product was stronger overall, but none caught the accounting nuance above without human-expert instruction. Reliability also varied — one evaluator model needed roughly 15 reruns to produce a usable Liquidation Valuation assessment, markedly less consistent than the other two.
- Test 2: Allow the agent to fetch a checkable fact when a needed figure is not available anywhere in the data room, and verify whether external data access changes the quality of its analysis vs. Test 1.
- Test 3: Allow the agent to fetch any open data on the internet (both objective and subjective data, e.g., forecasts, analysis, or opinions), and verify whether this changes the quality of its analysis vs. Test 2.
- Test 4: In Run E, allow Claude Cowork to fetch any open data, without access to the pre-built data room. The prompt would be "Analyze distressed investment opportunity for Beyond Meat and propose a strategy." None of Foresight's prompts will be used. Verify whether this changes the quality of its analysis vs. Run D-1.