← All topics

futuresearch

2 captures, most recent first.

Saved image — no attribution recorded

— saved image

The FutureSearch tools have disconnected following the server restart — none of the `mcp__futuresearch__*` tools are currently available in this session. This is on the connection/environment side, not something I can fix from here. **Good news:** the forecast itself runs entirely server-side on FutureSearch's platform, independent of this chat connection, so it should still be running (or may have already finished) regardless of this tool disconnection. **What I'd suggest:** - Session ID: `8ff4b059-7a6a-425d-aebb-63c95fc9ac87` - Task ID: `f8a77b1b-75f4-4bc1-b597-d1e9b9ae84c5` You can check the FutureSearch dashboard directly for this session to see if it's completed. Alternatively, if you can reconnect/refresh the FutureSearch tool integration on your end, let me know and I'll resume polling immediately.
Note from Claude Sonnet 5

Continuation of the same AI agent chat log as the prior screenshot: the agent reports its FutureSearch MCP tools disconnected after a server restart, gives session/task IDs again, and suggests checking the FutureSearch dashboard directly or reconnecting the tool integration.

ai agentsmcp toolsfuturesearchforecasting task

Dan Schwarz @dschwarz26

@dschwarz26 (Dan Schwarz) — 2h First Claude Fable forecasting evals are up. Fable is the best single-agent researcher as judged by predicting 1k near-term business, science, and technology outcomes, but not by a statistically significant amount. We unfortunately didn't run Fable-xhigh or Fable-max in time. (Each of those runs would have cost thousands of dollars, a single agent driven by those models can cost a few dollars.) Obviously we also can't use Fable in FutureSearch's best forecaster, the one you get in the app. But we are building around it, for when it becomes available. We'll update evals.futuresearch.ai as we dig through the agent traces. I'm curious to see if Fable's strategic reasoning failure modes match Opus, e.g. failing to judge political incentives as well as humans. [Embedded image: table titled "BTF-3 Leaderboard, Evaluated: June 2026", subtitle "All scores are on the Brier scale; LOWER IS BETTER, and the best score in each column is bolded." Columns: AGENT | POOLED SCORE (n=1,007) | BINARY (Brier, n=759) | NUMERIC (RPS, n=248) 1. FutureSearch SOTA* — 0.116 [0.106-0.127] | 0.114 [0.100-0.128] | 0.123 [0.110-0.137] 2. Claude Fable 5 (high) — 0.126 [0.114-0.137] | 0.124 [0.109-0.140] | 0.130 [0.117-0.144] 3. Claude Opus 4.8 (xhigh) — 0.127 [0.116-0.138] | 0.126 [0.112-0.141] | 0.130 [0.117-0.143] 4. GPT-5.5 (agent SDK)‡ — 0.127 [0.118-0.136] | 0.129 [0.118-0.140] | 0.122 [0.109-0.135] (bolded, best numeric) 5. Claude Opus 4.8 (high) — 0.134 [0.123-0.145] | 0.128 [0.114-0.143] | 0.152 [0.138-0.165] (table cut off at bottom, more rows likely below)]
Note from Claude Sonnet 5

Tweet with an embedded benchmark leaderboard table comparing forecasting accuracy of Claude Fable, Claude Opus 4.8, and GPT-5.5 variants.

claude fableforecastingai benchmarksfuturesearchclaude opus