TERMINAL-BENCH //
Stress-Test & Train Financial Agents in Real CLI Environments
Terminal-Bench+ Analytics, the Economics and Finance extension for agentic coding, captures professional financial and enterprise workflows – from portfolio analysis and return computation to volatility modeling and performance reporting – in a live terminal environment.
Developed by Snorkel’s AI Data Research Lab, this dataset provides high-signal, expert-created tasks built to challenge the reasoning limits of frontier models across complex finance workflows.
Agentic task categories include
Asset price & return evaluation
Computes returns, volatility, and risk-adjusted metrics for individual traded instruments.Structured financial records analysis
Extracts and aggregates metrics from structured records like income statements and balance sheets.Analytical reporting
Transforms raw tabular inputs into clean, aggregated datasets ready for downstream analysis.Cross-entity performance analysis
Evaluates, compares, and ranks multiple entities using consistent derived metrics.
Terminal-Bench+ is intentionally calibrated to stress state-of-the-art models
Built for Frontier model evaluation.
- Tiered difficulty from Core to Frontier
- <40% accuracy on Frontier tasks across leading models
- Designed for RL training, benchmarking, and deployment validation
If your agent succeeds here, it performs in production.