Terminal-Bench-Science Is Testing Agents on Actual Scientific Work
Terminal-Bench-Science 0.1 grades agents on researcher-contributed workflows, not textbook Q&A—and the best system still only clears about 30% of tasks.
Inspired by terminal-bench-science.ai · Steven Dillmann