In Test 1, Foresight v1 — working from a closed data room — outperformed senior human experts, and in Test 2, letting it fetch only the necessary missing facts improved its output further. In Test 3, once Foresight fetches open data entirely on its own, without a data room at all, output quality drops significantly. This run scores roughly 44% lower on average than the prior two runs, though it is still 7.5% higher than senior human experts.
- Run D's self-collected sources were all relevant to the requested opportunity, with no noise — but they relied heavily on news coverage while still capturing the primary financial numbers needed for model building.
- None of the runs showed AI hallucination, because the prompts already built in safeguards: chain-of-thought reasoning, a self-assessment pass, and mandatory traceability of every claim back to a tagged source.
- Two post-Q3 2025 figures were missed because Run D worked from raw, un-tagged sources. A source set heavy in press releases makes footnote-level figures easy to miss, and the agent captured fragmented data it couldn't assemble into a whole figure.
- Open-data fetching required a heavier setup — Claude Cowork to store and organize the sourced material, plus a separate tool for calculation. That meant ~1.5 hours (vs. 1 hour) and 1,278,136 output tokens, costing about $32 (vs. $4–$5) at Claude Opus 4.8's list rate of $25 per million output tokens.
Takeaway: Open-data fetching is still directionally reliable despite the drop in accuracy — feasible and cost-effective as a first-pass screen across a large set of opportunities, at a fraction of the cost of senior human experts. The prompt still needs refinement to avoid missing key data already sitting in its own sources. For any executive deciding whether unrestricted open-data access is worth enabling, this is a concrete answer: it still beats senior human experts, and it opens the door to screening opportunities directly from real-time market signals.
This test addresses a question about AI's autonomous judgment over which data to trust: when an agent collects data entirely from open sources on its own, rather than from a closed data room, does the financial analysis drift — and if so, by how much, and at what cost?
The case remains Beyond Meat, Inc., the same distressed-investing opportunity analyzed in Test 1 and Test 2. The comparison is again Foresight against itself — now run under unrestricted, open-data fetching rules on an otherwise identical task, subject to the same Dec 8, 2025 cutoff used in the earlier tests.
- Completed by the four-person Stern EMBA team.
- Reused in full from Test 1: the Foresight output built from the closed data room, no external fetch.
- Reused in full from Test 2: the closed data room plus permission to fetch specific missing facts only.
- ~1 hours, 166K output tokens (~$5).
- A fresh Foresight run on the identical task, with no data room at all and full ability to look up anything from the internet (same Dec 8, 2025 cutoff).
- Run in Claude Cowork so each lookup's full context could be saved and organized.
- ~1.5 hours, 1,278K output tokens (~$32).
On-topic sources, but press-release-heavy
Run D's self-collected set was all on-topic for distressed investing and included the primary financial filing needed for modeling. But it skewed to news: ten of twelve items were company-specific news against a single financial statement, where the closed data room was anchored by SEC filings.
No AI hallucination
Even though Run D collected its own data and built the model from it, both evaluators explicitly checked and confirmed that sampled figures traced back to real documents, with no fabricated sources. The prompts' built-in safeguards — explicit reasoning, a self-assessment pass, and mandatory source traceability — held.
Two post-quarter figures never made it into the model
Run D missed Beyond Meat's $148.7M at-the-market (ATM) equity raise in the 13-week cash flow, and used pre-exchange debt instead of the post-exchange figure in the liquidation valuation. Both emerged after Q3 2025. One was dropped at data capture — a cash-positive raise ran counter to the distress thesis, and no prompt forced the agent to reconcile it; the other was captured only as fragments the agent never summed into a single number.
Open-data access lost to both prior runs
Run D-1 slightly beat Run A, where its cleaner model architecture outweighed its shortfall on the two missed figures. Against Run B and Run C — both anchored by the data room — the open-data run was judged weaker in every comparison.
"Run D-1 delivers a complete, auditable, multi-scenario model with graded critical assumptions, explicit working-capital stress, self-validation checks, and a clear decision handoff." — Grok, 13-Week Cash Flow (Run D-1 vs. Run A)
"Run A better captures the post-Q3 financing/capital-structure change and produces security-level recoveries, while Run D-1 leaves its headline denominator at pre-exchange funded debt and omits the documented ATM inflow." — ChatGPT, Liquidation Valuation (Run D-1 vs. Run A)
"Run B is stronger because it includes the documented ATM proceeds and therefore gives the more defensible direction for near-term cash. Run D-1 has cleaner PIK debt-service treatment and more direct primary-source labeling, but omitting the $148.7 million ATM materially understates cash and overwhelms those advantages." — ChatGPT, 13-Week Cash Flow (Run D-1 vs. Run B)
"Run C is ready to use as-is for hand-off to the Priority-3 debt waterfall. Run D-1 should not be used for any Dec-8 decision without a complete rebuild of the valuation-date balance sheet and the post-exchange claim stack." — Grok, Liquidation Valuation (Run D-1 vs. Run C)
Root-cause analysis still needs senior human experts
AI cannot investigate the root cause of missing data on its own; it stays at the surface and still depends on senior human experts to challenge its answers. But Cowork — where each run produces 20+ generated files — speeds up the file-by-file comparison, so an expert can pinpoint the root cause within about an hour.
Open-data fetch is still unstable
A second collection round gathered less than the first — pulling only the Q3 2025 earnings-call transcript rather than the 10-Q — which shifts the analytical plan and how the model is built. Stabilizing the fetch is targeted for development under Foresight v2.
- Test 4: In Run E, allow Claude Cowork to fetch any open data, without access to the pre-built data room. The prompt would be "Analyze distressed investment opportunity for Beyond Meat and propose a strategy." None of Foresight's prompts will be used. Verify whether this changes the quality of its analysis vs. Run D-1.