Dean W. Ball @deanwball
Tweet: "in case you needed any more evidence that the reasoning/reinforcement learning approach is not limited to math and code (from the deep research system card)". Embedded quote from a system card describing how the "Deep Research" model was trained via reinforcement learning on browsing datasets to search, click, scroll, use a python sandbox for calculations and plotting, and synthesize many websites into reports.
Note from Claude Sonnet 5
Dean Ball highlights a system-card passage as evidence that RL-based reasoning training generalizes beyond narrow math and code domains to open-ended web research tasks.