Plots
Interactive charts generated from my AI-forecasting data pipeline. Dates are when the underlying data was last refreshed.
- BECI: AI model capability Every model's BECI capability score, plotted against release date by default and against measured inference cost, task latency or tokens spent at the flip of a switch. Colour by organization, country or open weights; filter to US or Chinese labs; drag the release-date cutoff; overlay benchmark difficulties (BEDI); switch to a full leaderboard; click through to any model's scorecard or any benchmark's fit page.
- Benchmarks Every benchmark's raw scores over time with the BECI logistic fit overlaid — searchable, with per-benchmark fit diagnostics.
- Models Pick up to four models and compare them across every benchmark they share — best score highlighted, per-model win tally, tag and test-set filters. A single model gets its full scorecard with ranks and fit residuals.
- Model size estimates Parameter-count estimates for frontier models from the model-sizes analysis.