# Trace - lecture_12

https://cs336.stanford.edu/lectures/?trace=lecture_12

lecture_12.py☀️⚪️🅴⬛⬅️➡️↖️↗️⤴️
1from edtrace import text, link, image
2from lecture_util import post_link
3from references import mmlu_2021
4
5def main():
6
7
8
9
10
11
12
13
14 what_is_good()
15
16 perplexity()
17 exam_benchmarks()
18 chat_benchmarks()
19 agentic_benchmarks()
20 pure_reasoning_benchmarks()
21 safety_benchmarks()
22
23 realism()
24 validity()
25 how_to_think_about_evaluation()
26
27
28
29
30
31
32
33def what_is_good():
34
35
36
37
38
39
40
41
42
43
44
45
[[Artificial Analysis]](https://artificialanalysis.ai/)
Artificial Analysis
46
![](https://cs336.stanford.edu/lectures/images/artificial-analysis.png)
47
48
49
![](https://cs336.stanford.edu/lectures/images/artificial-analysis-cost.png)
50
51
52
[[Arena AI (formerly Chatbot Arena)]](https://arena.ai/leaderboard)
Arena AI (formerly Chatbot Arena)
53
![](https://cs336.stanford.edu/lectures/images/lmarena-leaderboard.png)
54
55
56
[[OpenRouter]](https://openrouter.ai/rankings)
OpenRouter
57
![](https://cs336.stanford.edu/lectures/images/openrouter.png)
58
59
60def perplexity():
61
62
63
64
65
66
67
68
69
70
71
72
73
[[Jozefowicz+ 2016]](https://arxiv.org/abs/1602.02410)
Exploring the Limits of Language Modeling
Rafal Jozefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, Yonghui Wu
2016-02-07
In this work we explore recent advances in Recurrent Neural Networks for large scale Language Modeling, a task central to language understanding. We extend current models to deal with two key challenges present in this task: corpora and vocabulary sizes, and complex, long term structure of language. We perform an exhaustive study on techniques such as character Convolutional Neural Networks or Long-Short Term Memory, on the One Billion Word Benchmark. Our best single model significantly improves state-of-the-art perplexity from 51.3 down to 30.0 (whilst reducing the number of parameters by a factor of 20), while an ensemble of models sets a new record by improving perplexity from 41.0 down to 23.7. We also release these models for the NLP and ML community to study and improve upon.
74
75
76
77
78
![](https://cs336.stanford.edu/lectures/images/gpt2-perplexity.png)
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
[[Paperno+ 2016]](https://arxiv.org/abs/1606.06031)
The LAMBADA dataset: Word prediction requiring a broad discourse context
Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Quan Ngoc Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, Raquel Fernández
2016-06-20
We introduce LAMBADA, a dataset to evaluate the capabilities of computational models for text understanding by means of a word prediction task. LAMBADA is a collection of narrative passages sharing the characteristic that human subjects are able to guess their last word if they are exposed to the whole passage, but not if they only see the last sentence preceding the target word. To succeed on LAMBADA, computational models cannot simply rely on local context, but must be able to keep track of information in the broader discourse. We show that LAMBADA exemplifies a wide range of linguistic phenomena, and that none of several state-of-the-art language models reaches accuracy above 1% on this novel benchmark. We thus propose LAMBADA as a challenging test set, meant to encourage the development of new models capable of genuine understanding of broad context in natural language text.
94
![](https://cs336.stanford.edu/lectures/images/lambada.png)
95
[[Zellers+ 2019]](https://arxiv.org/pdf/1905.07830)
HellaSwag: Can a Machine Really Finish Your Sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, Yejin Choi
2019-05-19
Recent work by Zellers et al. (2018) introduced a new task of commonsense natural language inference: given an event description such as "A woman sits at a piano," a machine must select the most likely followup: "She sets her fingers on the keys." With the introduction of BERT, near human-level performance was reached. Does this mean that machines can perform human level commonsense inference? In this paper, we show that commonsense inference still proves difficult for even state-of-the-art models, by presenting HellaSwag, a new challenge dataset. Though its questions are trivial for humans (>95% accuracy), state-of-the-art models struggle (<48%). We achieve this via Adversarial Filtering (AF), a data collection paradigm wherein a series of discriminators iteratively select an adversarial set of machine-generated wrong answers. AF proves to be surprisingly robust. The key insight is to scale up the length and complexity of the dataset examples towards a critical 'Goldilocks' zone wherein generated text is ridiculous to humans, yet often misclassified by state-of-the-art models. Our construction of HellaSwag, and its resulting difficulty, sheds light on the inner workings of deep pretrained models. More broadly, it suggests a new path forward for NLP research, in which benchmarks co-evolve with the evolving state-of-the-art in an adversarial way, so as to present ever-harder challenges.
96
![](https://cs336.stanford.edu/lectures/images/hellaswag.png)
97
98
99
100
101
102
103
104
105
106
107
108def exam_benchmarks():
109
110
111
112
113
[[Hendrycks+ 2020]](https://arxiv.org/pdf/2009.03300.pdf)
Measuring Massive Multitask Language Understanding
[Berkeley] Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, Jacob Steinhardt
2020-09-07
We propose a new test to measure a text model's multitask accuracy. The test covers 57 tasks including elementary mathematics, US history, computer science, law, and more. To attain high accuracy on this test, models must possess extensive world knowledge and problem solving ability. We find that while most recent models have near random-chance accuracy, the very largest GPT-3 model improves over random chance by almost 20 percentage points on average. However, on every one of the 57 tasks, the best models still need substantial improvements before they can reach expert-level accuracy. Models also have lopsided performance and frequently do not know when they are wrong. Worse, they still have near-random accuracy on some socially important subjects such as morality and law. By comprehensively evaluating the breadth and depth of a model's academic and professional understanding, our test can be used to analyze models across many tasks and to identify important shortcomings.
57 subjects, multiple-choice
114
115
116
117
118
![](https://cs336.stanford.edu/lectures/images/mmlu.png)
119
[[https://llm-stats.com/benchmarks/mmlu]](https://llm-stats.com/benchmarks/mmlu)
120
[[HELM MMLU for visualizing predictions]](https://crfm.stanford.edu/helm/mmlu/latest/)
HELM MMLU for visualizing predictions
121
122
[[Wang+ 2024]](https://arxiv.org/abs/2406.01574)
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, Abhranil Chandra ... (7 more) ... Kai Wang, Alex Zhuang, Rongqi Fan, Xiang Yue, Wenhu Chen
2024-06-03
In the age of large-scale language models, benchmarks like the Massive Multitask Language Understanding (MMLU) have been pivotal in pushing the boundaries of what AI can achieve in language comprehension and reasoning across diverse domains. However, as models continue to improve, their performance on these benchmarks has begun to plateau, making it increasingly difficult to discern differences in model capabilities. This paper introduces MMLU-Pro, an enhanced dataset designed to extend the mostly knowledge-driven MMLU benchmark by integrating more challenging, reasoning-focused questions and expanding the choice set from four to ten options. Additionally, MMLU-Pro eliminates the trivial and noisy questions in MMLU. Our experimental results show that MMLU-Pro not only raises the challenge, causing a significant drop in accuracy by 16% to 33% compared to MMLU but also demonstrates greater stability under varying prompts. With 24 different prompt styles tested, the sensitivity of model scores to prompt variations decreased from 4-5% in MMLU to just 2% in MMLU-Pro. Additionally, we found that models utilizing Chain of Thought (CoT) reasoning achieved better performance on MMLU-Pro compared to direct answering, which is in stark contrast to the findings on the original MMLU, indicating that MMLU-Pro includes more complex reasoning questions. Our assessments confirm that MMLU-Pro is a more discriminative benchmark to better track progress in the field.
123
124
125
126
127
![](https://cs336.stanford.edu/lectures/images/mmlu-pro.png)
128
[[https://llm-stats.com/benchmarks/mmlu-pro]](https://llm-stats.com/benchmarks/mmlu-pro)
129
[[HELM MMLU-Pro for visualizing predictions]](https://crfm.stanford.edu/helm/capabilities/latest/#/leaderboard/mmlu_pro)
HELM MMLU-Pro for visualizing predictions
130
131
[[Rein+ 2023]](https://arxiv.org/abs/2311.12022)
GPQA: A Graduate-Level Google-Proof Q&A Benchmark
David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, Samuel R. Bowman
2023-11-20
We present GPQA, a challenging dataset of 448 multiple-choice questions written by domain experts in biology, physics, and chemistry. We ensure that the questions are high-quality and extremely difficult: experts who have or are pursuing PhDs in the corresponding domains reach 65% accuracy (74% when discounting clear mistakes the experts identified in retrospect), while highly skilled non-expert validators only reach 34% accuracy, despite spending on average over 30 minutes with unrestricted access to the web (i.e., the questions are "Google-proof"). The questions are also difficult for state-of-the-art AI systems, with our strongest GPT-4 based baseline achieving 39% accuracy. If we are to use future AI systems to help us answer very hard questions, for example, when developing new scientific knowledge, we need to develop scalable oversight methods that enable humans to supervise their outputs, which may be difficult even if the supervisors are themselves skilled and knowledgeable. The difficulty of GPQA both for skilled non-experts and frontier AI systems should enable realistic scalable oversight experiments, which we hope can help devise ways for human experts to reliably get truthful information from AI systems that surpass human capabilities.
132
133
![](https://cs336.stanford.edu/lectures/images/gpqa.png)
134
135
136
137
[[https://llm-stats.com/benchmarks/gpqa]](https://llm-stats.com/benchmarks/gpqa)
138
[[HELM GPQA for visualizing predictions]](https://crfm.stanford.edu/helm/capabilities/latest/#/leaderboard/gpqa)
HELM GPQA for visualizing predictions
139
140
[[Phan+ 2025]](https://arxiv.org/abs/2501.14249)
Humanity's Last Exam
Long Phan, Alice Gatti, Ziwen Han, Nathaniel Li, Josephina Hu ... (1110 more) ... Ido Akov, Artem Lukoianov, Summer Yue, Alexandr Wang, Dan Hendrycks
2025-01-24
Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achieve over 90\% accuracy on popular benchmarks like MMLU, limiting informed measurement of state-of-the-art LLM capabilities. In response, we introduce Humanity's Last Exam (HLE), a multi-modal benchmark at the frontier of human knowledge, designed to be the final closed-ended academic benchmark of its kind with broad subject coverage. HLE consists of 2,500 questions across dozens of subjects, including mathematics, humanities, and the natural sciences. HLE is developed globally by subject-matter experts and consists of multiple-choice and short-answer questions suitable for automated grading. Each question has a known solution that is unambiguous and easily verifiable, but cannot be quickly answered via internet retrieval. State-of-the-art LLMs demonstrate low accuracy and calibration on HLE, highlighting a significant gap between current LLM capabilities and the expert human frontier on closed-ended academic questions. To inform research and policymaking upon a clear understanding of model capabilities, we publicly release HLE at https://lastexam.ai.
141
142
![](https://cs336.stanford.edu/lectures/images/hle-examples.png)
143
144
145
![](https://cs336.stanford.edu/lectures/images/hle-pipeline.png)
146
![](https://cs336.stanford.edu/lectures/images/hle-results.png)
147
[[https://llm-stats.com/benchmarks/hle]](https://llm-stats.com/benchmarks/hle)
148
149
150
151
152
153
154
155def chat_benchmarks():
156
157
158
159
160
161
162
163
164
165
[[Chiang+ 2024]](https://arxiv.org/abs/2403.04132)
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li ... (1 more) ... Hao Zhang, Banghua Zhu, Michael Jordan, Joseph E. Gonzalez, Ion Stoica
2024-03-07
Large Language Models (LLMs) have unlocked new capabilities and applications; however, evaluating the alignment with human preferences still poses significant challenges. To address this issue, we introduce Chatbot Arena, an open platform for evaluating LLMs based on human preferences. Our methodology employs a pairwise comparison approach and leverages input from a diverse user base through crowdsourcing. The platform has been operational for several months, amassing over 240K votes. This paper describes the platform, analyzes the data we have collected so far, and explains the tried-and-true statistical methods we are using for efficient and accurate evaluation and ranking of models. We confirm that the crowdsourced questions are sufficiently diverse and discriminating and that the crowdsourced human votes are in good agreement with those of expert raters. These analyses collectively establish a robust foundation for the credibility of Chatbot Arena. Because of its unique value and openness, Chatbot Arena has emerged as one of the most referenced LLM leaderboards, widely cited by leading LLM developers and companies. Our demo is publicly available at \url{https://chat.lmsys.org}.
166
167
168
169
170
![](https://cs336.stanford.edu/lectures/images/arena-beets.png)
171
172
173
174
[[Arena AI (formerly Chatbot Arena)]](https://arena.ai/leaderboard)
Arena AI (formerly Chatbot Arena)
175
![](https://cs336.stanford.edu/lectures/images/lmarena-leaderboard.png)
176
177
178
179
180
181
182
183
184
[[leaderboard]](https://tatsu-lab.github.io/alpaca_eval/)
leaderboard
185
186
187
188
[[Dubois+ 2024]](https://arxiv.org/pdf/2404.04475)
Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
Yann Dubois, Balázs Galambosi, Percy Liang, Tatsunori B. Hashimoto
2024-04-06
LLM-based auto-annotators have become a key component of the LLM development process due to their cost-effectiveness and scalability compared to human-based evaluation. However, these auto-annotators can introduce biases that are hard to remove. Even simple, known confounders such as preference for longer outputs remain in existing automated evaluation metrics. We propose a simple regression analysis approach for controlling biases in auto-evaluations. As a real case study, we focus on reducing the length bias of AlpacaEval, a fast and affordable benchmark for instruction-tuned LLMs that uses LLMs to estimate response quality. Despite being highly correlated with human preferences, AlpacaEval is known to favor models that generate longer outputs. We introduce a length-controlled AlpacaEval that aims to answer the counterfactual question: "What would the preference be if the model's and baseline's output had the same length?" To achieve this, we first fit a generalized linear model to predict the biased auto-annotator's preferences based on the mediators we want to control for (length difference) and other relevant features. We then obtain length-controlled preferences by predicting preferences while conditioning the GLM with a zero difference in lengths. Length-controlling not only improves the robustness of the metric to manipulations in model verbosity, but we also find that it increases the Spearman correlation with LMSYS Chatbot Arena from 0.94 to 0.98.
189
190
191
![](https://cs336.stanford.edu/lectures/var/files/image-434a1510a7ed21d5355814149a9490c4-https_github_com_tatsu-lab_alpaca_eval_raw_main_figures_chat_correlations_no_ae_png)
192
![](https://cs336.stanford.edu/lectures/images/alpacaeval-leaderboard.png)
193
194
[[Lin+ 2024]](https://arxiv.org/pdf/2406.04770)
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
Bill Yuchen Lin, Yuntian Deng, Khyathi Chandu, Faeze Brahman, Abhilasha Ravichander, Valentina Pyatkin, Nouha Dziri, Ronan Le Bras, Yejin Choi
2024-06-07
We introduce WildBench, an automated evaluation framework designed to benchmark large language models (LLMs) using challenging, real-world user queries. WildBench consists of 1,024 tasks carefully selected from over one million human-chatbot conversation logs. For automated evaluation with WildBench, we have developed two metrics, WB-Reward and WB-Score, which are computable using advanced LLMs such as GPT-4-turbo. WildBench evaluation uses task-specific checklists to evaluate model outputs systematically and provides structured explanations that justify the scores and comparisons, resulting in more reliable and interpretable automatic judgments. WB-Reward employs fine-grained pairwise comparisons between model responses, generating five potential outcomes: much better, slightly better, slightly worse, much worse, or a tie. Unlike previous evaluations that employed a single baseline model, we selected three baseline models at varying performance levels to ensure a comprehensive pairwise evaluation. Additionally, we propose a simple method to mitigate length bias, by converting outcomes of ``slightly better/worse'' to ``tie'' if the winner response exceeds the loser one by more than $K$ characters. WB-Score evaluates the quality of model outputs individually, making it a fast and cost-efficient evaluation metric. WildBench results demonstrate a strong correlation with the human-voted Elo ratings from Chatbot Arena on hard tasks. Specifically, WB-Reward achieves a Pearson correlation of 0.98 with top-ranking models. Additionally, WB-Score reaches 0.95, surpassing both ArenaHard's 0.91 and AlpacaEval2.0's 0.89 for length-controlled win rates, as well as the 0.87 for regular win rates.
195
196
197
198
![](https://cs336.stanford.edu/lectures/images/wildbench.png)
199
[[HELM WildBench for visualizing predictions]](https://crfm.stanford.edu/helm/capabilities/latest/#/leaderboard/wildbench)
HELM WildBench for visualizing predictions
200
201
202
203
204
205
206
207
208def agentic_benchmarks():
209
210
211
212
213
214
215
216
[[Jimenez+ 2023]](https://arxiv.org/abs/2310.06770)
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, Karthik Narasimhan
2023-10-10
Language models have outpaced our ability to evaluate them effectively, but for their future development it is essential to study the frontier of their capabilities. We find real-world software engineering to be a rich, sustainable, and challenging testbed for evaluating the next generation of language models. To this end, we introduce SWE-bench, an evaluation framework consisting of $2,294$ software engineering problems drawn from real GitHub issues and corresponding pull requests across $12$ popular Python repositories. Given a codebase along with a description of an issue to be resolved, a language model is tasked with editing the codebase to address the issue. Resolving issues in SWE-bench frequently requires understanding and coordinating changes across multiple functions, classes, and even files simultaneously, calling for models to interact with execution environments, process extremely long contexts and perform complex reasoning that goes far beyond traditional code generation tasks. Our evaluations show that both state-of-the-art proprietary models and our fine-tuned model SWE-Llama can resolve only the simplest issues. The best-performing model, Claude 2, is able to solve a mere $1.96$% of the issues. Advances on SWE-bench represent steps towards LMs that are more practical, intelligent, and autonomous.
217
218
219
220
![](https://cs336.stanford.edu/lectures/images/swebench.png)
221
[[https://llm-stats.com/benchmarks/swe-bench-verified]](https://llm-stats.com/benchmarks/swe-bench-verified)
222
223
[[Merrill+ 2026]](https://arxiv.org/abs/2601.11868)
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Mike A. Merrill, Alexander G. Shaw, Nicholas Carlini, Boxuan Li, Harsh Raj ... (75 more) ... Ryan Marten, Yixin Wang, Alex Dimakis, Andy Konwinski, Ludwig Schmidt
2026-01-17
AI agents may soon become capable of autonomously completing valuable, long-horizon tasks in diverse domains. Current benchmarks either do not measure real-world tasks, or are not sufficiently difficult to meaningfully measure frontier models. To this end, we present Terminal-Bench 2.0: a carefully curated hard benchmark composed of 89 tasks in computer terminal environments inspired by problems from real workflows. Each task features a unique environment, human-written solution, and comprehensive tests for verification. We show that frontier models and agents score less than 65\% on the benchmark and conduct an error analysis to identify areas for model and agent improvement. We publish the dataset and evaluation harness to assist developers and researchers in future work at https://www.tbench.ai/ .
[[website]](https://www.tbench.ai/)
website
224
![](https://cs336.stanford.edu/lectures/images/terminal-bench.png)
225
226
227
![](https://cs336.stanford.edu/lectures/images/terminal-bench-human-time.png)
228
![](https://cs336.stanford.edu/lectures/images/terminal-bench-results.png)
229
[[https://llm-stats.com/benchmarks/terminal-bench]](https://llm-stats.com/benchmarks/terminal-bench)
230
231
[[Zhang+ 2024]](https://arxiv.org/abs/2408.08926)
Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Andy K. Zhang, Neil Perry, Riya Dulepet, Joey Ji, Celeste Menders ... (17 more) ... Kenny Osele, Gautham Raghupathi, Dan Boneh, Daniel E. Ho, Percy Liang
2024-08-15
Language Model (LM) agents for cybersecurity that are capable of autonomously identifying vulnerabilities and executing exploits have potential to cause real-world impact. Policymakers, model providers, and researchers in the AI and cybersecurity communities are interested in quantifying the capabilities of such agents to help mitigate cyberrisk and investigate opportunities for penetration testing. Toward that end, we introduce Cybench, a framework for specifying cybersecurity tasks and evaluating agents on those tasks. We include 40 professional-level Capture the Flag (CTF) tasks from 4 distinct CTF competitions, chosen to be recent, meaningful, and spanning a wide range of difficulties. Each task includes its own description, starter files, and is initialized in an environment where an agent can execute commands and observe outputs. Since many tasks are beyond the capabilities of existing LM agents, we introduce subtasks for each task, which break down a task into intermediary steps for a more detailed evaluation. To evaluate agent capabilities, we construct a cybersecurity agent and evaluate 8 models: GPT-4o, OpenAI o1-preview, Claude 3 Opus, Claude 3.5 Sonnet, Mixtral 8x22b Instruct, Gemini 1.5 Pro, Llama 3 70B Chat, and Llama 3.1 405B Instruct. For the top performing models (GPT-4o and Claude 3.5 Sonnet), we further investigate performance across 4 agent scaffolds (structed bash, action-only, pseudoterminal, and web search). Without subtask guidance, agents leveraging Claude 3.5 Sonnet, GPT-4o, OpenAI o1-preview, and Claude 3 Opus successfully solved complete tasks that took human teams up to 11 minutes to solve. In comparison, the most difficult task took human teams 24 hours and 54 minutes to solve. All code and data are publicly available at https://cybench.github.io.
232
![](https://cs336.stanford.edu/lectures/images/cybench.png)
233
234
235
![](https://cs336.stanford.edu/lectures/images/cybench-agent.png)
236
![](https://cs336.stanford.edu/lectures/images/cybench-results.png)
237
[[https://llm-stats.com/benchmarks/cybench]](https://llm-stats.com/benchmarks/cybench)
238
239
[[Chan+ 2024]](https://arxiv.org/abs/2410.07095)
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Jun Shern Chan, Neil Chowdhury, Oliver Jaffe, James Aung, Dane Sherburn ... (2 more) ... Kevin Liu, Leon Maksin, Tejal Patwardhan, Lilian Weng, Aleksander Mądry
2024-10-09
We introduce MLE-bench, a benchmark for measuring how well AI agents perform at machine learning engineering. To this end, we curate 75 ML engineering-related competitions from Kaggle, creating a diverse set of challenging tasks that test real-world ML engineering skills such as training models, preparing datasets, and running experiments. We establish human baselines for each competition using Kaggle's publicly available leaderboards. We use open-source agent scaffolds to evaluate several frontier language models on our benchmark, finding that the best-performing setup--OpenAI's o1-preview with AIDE scaffolding--achieves at least the level of a Kaggle bronze medal in 16.9% of competitions. In addition to our main results, we investigate various forms of resource scaling for AI agents and the impact of contamination from pre-training. We open-source our benchmark code (github.com/openai/mle-bench/) to facilitate future research in understanding the ML engineering capabilities of AI agents.
240
241
![](https://cs336.stanford.edu/lectures/images/mlebench.png)
242
![](https://cs336.stanford.edu/lectures/images/mlebench-results.png)
243
244
[[post]](https://www.philschmid.de/agents-2.0-deep-agents)
post
245
![](https://cs336.stanford.edu/lectures/var/files/image-155d1eb10710df090449bf401822dd7e-https_www_philschmid_de_static_blog_agents-2_0-deep-agents_overview_png)
246
247
248
249
250
251
252
253
254
255
256
257def pure_reasoning_benchmarks():
258
259
260
261
262
[[website]](https://arcprize.org/arc-agi)
website
263
264
265
266
267
![](https://cs336.stanford.edu/lectures/var/files/image-d1a33e9159cfdb77197551bbbecc6a76-https_arcprize_org_media_images_arc-task-grids_jpg)
268
269
270
![](https://cs336.stanford.edu/lectures/var/files/image-a0338c9fb72d1163cfc8ac66fea4e4ed-https_arcprize_org_media_images_blog_arc-agi-2-unsolved-1_png)
271
272
![](https://cs336.stanford.edu/lectures/images/arc-agi-results.png)
273
274
275
276
[[post]](https://arcprize.org/media/ARC_AGI_3_Technical_Report.pdf)
post
277
![](https://cs336.stanford.edu/lectures/images/arc-agi-3.png)
278
![](https://cs336.stanford.edu/lectures/images/arc-agi-3-results.png)
279
280
281
282
283
284
285
286def safety_benchmarks():
287
![](https://cs336.stanford.edu/lectures/var/files/image-a375cd28c372458baf4135c081a1ce8b-https_www_team-bhp_com_forum_attachments_road-safety_2173645d1625144681-will-crash-test-rating-change-if-higher-variant-chosen-images-30_jpeg)
288
289
290
[[Mazeika+ 2024]](https://arxiv.org/abs/2402.04249)
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang ... (2 more) ... Nathaniel Li, Steven Basart, Bo Li, David Forsyth, Dan Hendrycks
2024-02-06
Automated red teaming holds substantial promise for uncovering and mitigating the risks associated with the malicious use of large language models (LLMs), yet the field lacks a standardized evaluation framework to rigorously assess new methods. To address this issue, we introduce HarmBench, a standardized evaluation framework for automated red teaming. We identify several desirable properties previously unaccounted for in red teaming evaluations and systematically design HarmBench to meet these criteria. Using HarmBench, we conduct a large-scale comparison of 18 red teaming methods and 33 target LLMs and defenses, yielding novel insights. We also introduce a highly efficient adversarial training method that greatly enhances LLM robustness across a wide range of attacks, demonstrating how HarmBench enables codevelopment of attacks and defenses. We open source HarmBench at https://github.com/centerforaisafety/HarmBench.
291
292
[[HarmBench on HELM]](https://crfm.stanford.edu/helm/safety/latest/#/leaderboard/harm_bench)
HarmBench on HELM
293
[[Example of safety failure]](https://crfm.stanford.edu/helm/safety/latest/#/runs/harm_bench:model=anthropic_claude-3-7-sonnet-20250219?instancesPage=4)
Example of safety failure
294
295
[[Zeng+ 2024]](https://arxiv.org/abs/2407.17436)
AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies
Yi Zeng, Yu Yang, Andy Zhou, Jeffrey Ziwei Tan, Yuheng Tu ... (2 more) ... Minzhou Pan, Ruoxi Jia, Dawn Song, Percy Liang, Bo Li
2024-07-11
Foundation models (FMs) provide societal benefits but also amplify risks. Governments, companies, and researchers have proposed regulatory frameworks, acceptable use policies, and safety benchmarks in response. However, existing public benchmarks often define safety categories based on previous literature, intuitions, or common sense, leading to disjointed sets of categories for risks specified in recent regulations and policies, which makes it challenging to evaluate and compare FMs across these benchmarks. To bridge this gap, we introduce AIR-Bench 2024, the first AI safety benchmark aligned with emerging government regulations and company policies, following the regulation-based safety categories grounded in our AI risks study, AIR 2024. AIR 2024 decomposes 8 government regulations and 16 company policies into a four-tiered safety taxonomy with 314 granular risk categories in the lowest tier. AIR-Bench 2024 contains 5,694 diverse prompts spanning these categories, with manual curation and human auditing to ensure quality. We evaluate leading language models on AIR-Bench 2024, uncovering insights into their alignment with specified safety concerns. By bridging the gap between public benchmarks and practical AI risks, AIR-Bench 2024 provides a foundation for assessing model safety across jurisdictions, fostering the development of safer and more responsible AI systems.
296
297
298
![](https://cs336.stanford.edu/lectures/var/files/image-5993188f3fa9dc78b85f9866fcee27ac-https_crfm_stanford_edu_helm_assets_air-overview-DpBbyagA_png)
299
[[HELM AIR-Bench]](https://crfm.stanford.edu/helm/air-bench/latest/#/leaderboard)
HELM AIR-Bench
300
301
302
303
[[Zou+ 2023]](https://arxiv.org/pdf/2307.15043)
Universal and Transferable Adversarial Attacks on Aligned Language Models
Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J. Zico Kolter, Matt Fredrikson
2023-07-27
Because "out-of-the-box" large language models are capable of generating a great deal of objectionable content, recent work has focused on aligning these models in an attempt to prevent undesirable generation. While there has been some success at circumventing these measures -- so-called "jailbreaks" against LLMs -- these attacks have required significant human ingenuity and are brittle in practice. In this paper, we propose a simple and effective attack method that causes aligned language models to generate objectionable behaviors. Specifically, our approach finds a suffix that, when attached to a wide range of queries for an LLM to produce objectionable content, aims to maximize the probability that the model produces an affirmative response (rather than refusing to answer). However, instead of relying on manual engineering, our approach automatically produces these adversarial suffixes by a combination of greedy and gradient-based search techniques, and also improves over past automatic prompt generation methods. Surprisingly, we find that the adversarial prompts generated by our approach are quite transferable, including to black-box, publicly released LLMs. Specifically, we train an adversarial attack suffix on multiple prompts (i.e., queries asking for many different types of objectionable content), as well as multiple models (in our case, Vicuna-7B and 13B). When doing so, the resulting attack suffix is able to induce objectionable content in the public interfaces to ChatGPT, Bard, and Claude, as well as open source LLMs such as LLaMA-2-Chat, Pythia, Falcon, and others. In total, this work significantly advances the state-of-the-art in adversarial attacks against aligned language models, raising important questions about how such systems can be prevented from producing objectionable information. Code is available at github.com/llm-attacks/llm-attacks.
304
305
![](https://cs336.stanford.edu/lectures/images/gcg-examples.png)
306
307
308
309
310
311
312
313
314def realism():
315
316
317
318
319
[[Patwardhan+ 2025]](https://arxiv.org/pdf/2510.04374)
GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
Tejal Patwardhan, Rachel Dias, Elizabeth Proehl, Grace Kim, Michele Wang ... (9 more) ... David Li, Michael Sharman, Alexandra Barr, Amelia Glaese, Jerry Tworek
2025-10-05
We introduce GDPval, a benchmark evaluating AI model capabilities on real-world economically valuable tasks. GDPval covers the majority of U.S. Bureau of Labor Statistics Work Activities for 44 occupations across the top 9 sectors contributing to U.S. GDP (Gross Domestic Product). Tasks are constructed from the representative work of industry professionals with an average of 14 years of experience. We find that frontier model performance on GDPval is improving roughly linearly over time, and that the current best frontier models are approaching industry experts in deliverable quality. We analyze the potential for frontier models, when paired with human oversight, to perform GDPval tasks cheaper and faster than unaided experts. We also demonstrate that increased reasoning effort, increased task context, and increased scaffolding improves model performance on GDPval. Finally, we open-source a gold subset of 220 tasks and provide a public automated grading service at evals.openai.com to facilitate future research in understanding real-world model capabilities.
320
321
322
![](https://cs336.stanford.edu/lectures/images/gdpval.png)
323
324
[[Bedi+ 2025]](https://arxiv.org/abs/2505.23802)
MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Suhana Bedi, Hejie Cui, Miguel Fuentes, Alyssa Unell, Michael Wornow ... (71 more) ... Matthew P. Lungren, Eric Horvitz, Percy Liang, Mike Pfeffer, Nigam H. Shah
2025-05-26
While large language models (LLMs) achieve near-perfect scores on medical licensing exams, these evaluations inadequately reflect the complexity and diversity of real-world clinical practice. We introduce MedHELM, an extensible evaluation framework for assessing LLM performance for medical tasks with three key contributions. First, a clinician-validated taxonomy spanning 5 categories, 22 subcategories, and 121 tasks developed with 29 clinicians. Second, a comprehensive benchmark suite comprising 35 benchmarks (17 existing, 18 newly formulated) providing complete coverage of all categories and subcategories in the taxonomy. Third, a systematic comparison of LLMs with improved evaluation methods (using an LLM-jury) and a cost-performance analysis. Evaluation of 9 frontier LLMs, using the 35 benchmarks, revealed significant performance variation. Advanced reasoning models (DeepSeek R1: 66% win-rate; o3-mini: 64% win-rate) demonstrated superior performance, though Claude 3.5 Sonnet achieved comparable results at 40% lower estimated computational cost. On a normalized accuracy scale (0-1), most models performed strongly in Clinical Note Generation (0.73-0.85) and Patient Communication & Education (0.78-0.83), moderately in Medical Research Assistance (0.65-0.75), and generally lower in Clinical Decision Support (0.56-0.72) and Administration & Workflow (0.53-0.63). Our LLM-jury evaluation method achieved good agreement with clinician ratings (ICC = 0.47), surpassing both average clinician-clinician agreement (ICC = 0.43) and automated baselines including ROUGE-L (0.36) and BERTScore-F1 (0.44). Claude 3.5 Sonnet achieved comparable performance to top models at lower estimated cost. These findings highlight the importance of real-world, task-specific evaluation for medical use of LLMs and provides an open source framework to enable this.
325
326
327
![](https://cs336.stanford.edu/lectures/var/files/image-93ff2615b50418e8fd4dd6f6435bdff1-https_crfm_stanford_edu_helm_assets_medhelm-overview-CND0EIsy_png)
328
[[MedHELM]](https://crfm.stanford.edu/helm/medhelm/latest/#/leaderboard)
MedHELM
329
330
[[Tamkin+ 2024]](https://arxiv.org/abs/2412.13678)
Clio: Privacy-Preserving Insights into Real-World AI Use
Alex Tamkin, Miles McCain, Kunal Handa, Esin Durmus, Liane Lovitt ... (11 more) ... Wes Mitchell, Shan Carter, Jack Clark, Jared Kaplan, Deep Ganguli
2024-12-18
How are AI assistants being used in the real world? While model providers in theory have a window into this impact via their users' data, both privacy concerns and practical challenges have made analyzing this data difficult. To address these issues, we present Clio (Claude insights and observations), a privacy-preserving platform that uses AI assistants themselves to analyze and surface aggregated usage patterns across millions of conversations, without the need for human reviewers to read raw conversations. We validate this can be done with a high degree of accuracy and privacy by conducting extensive evaluations. We demonstrate Clio's usefulness in two broad ways. First, we share insights about how models are being used in the real world from one million Claude.ai Free and Pro conversations, ranging from providing advice on hairstyles to providing guidance on Git operations and concepts. We also identify the most common high-level use cases on Claude.ai (coding, writing, and research tasks) as well as patterns that differ across languages (e.g., conversations in Japanese discuss elder care and aging populations at higher-than-typical rates). Second, we use Clio to make our systems safer by identifying coordinated attempts to abuse our systems, monitoring for unknown unknowns during critical periods like launches of new capabilities or major world events, and improving our existing monitoring systems. We also discuss the limitations of our approach, as well as risks and ethical concerns. By enabling analysis of real-world AI usage, Clio provides a scalable platform for empirically grounded AI safety and governance.
331
332
333
![](https://cs336.stanford.edu/lectures/images/clio-table4.png)
334
335
336
337
338def validity():
339
340
341
342
343
344
345
346
347
[[Oren+ 2023]](https://arxiv.org/pdf/2310.17623)
Proving Test Set Contamination in Black Box Language Models
Yonatan Oren, Nicole Meister, Niladri Chatterji, Faisal Ladhak, Tatsunori B. Hashimoto
2023-10-26
Large language models are trained on vast amounts of internet data, prompting concerns and speculation that they have memorized public benchmarks. Going from speculation to proof of contamination is challenging, as the pretraining data used by proprietary models are often not publicly accessible. We show that it is possible to provide provable guarantees of test set contamination in language models without access to pretraining data or model weights. Our approach leverages the fact that when there is no data contamination, all orderings of an exchangeable benchmark should be equally likely. In contrast, the tendency for language models to memorize example order means that a contaminated language model will find certain canonical orderings to be much more likely than others. Our test flags potential contamination whenever the likelihood of a canonically ordered benchmark dataset is significantly higher than the likelihood after shuffling the examples. We demonstrate that our procedure is sensitive enough to reliably prove test set contamination in challenging situations, including models as small as 1.4 billion parameters, on small test sets of only 1000 examples, and datasets that appear only a few times in the pretraining corpus. Using our test, we audit five popular publicly accessible language models for test set contamination and find little evidence for pervasive contamination.
348
![](https://cs336.stanford.edu/lectures/images/contamination-exchangeability.png)
349
350
351
[[Zhang+ 2024]](https://arxiv.org/abs/2410.08385)
Language model developers should report train-test overlap
Andy K Zhang, Kevin Klyman, Yifan Mai, Yoav Levine, Yian Zhang, Rishi Bommasani, Percy Liang
2024-10-10
Language models are extensively evaluated, but correctly interpreting evaluation results requires knowledge of train-test overlap which refers to the extent to which the language model is trained on the very data it is being tested on. The public currently lacks adequate information about train-test overlap: most models have no public train-test overlap statistics, and third parties cannot directly measure train-test overlap since they do not have access to the training data. To make this clear, we document the practices of 30 model developers, finding that just 9 developers report train-test overlap: 4 developers release training data under open-source licenses, enabling the community to directly measure train-test overlap, and 5 developers publish their train-test overlap methodology and statistics. By engaging with language model developers, we provide novel information about train-test overlap for three additional developers. Overall, we take the position that language model developers should publish train-test overlap statistics and/or training data whenever they report evaluation results on public test sets. We hope our work increases transparency into train-test overlap to increase the community-wide trust in model evaluations.
352
353
354
355
356
357
358
359
360
361
362
363
[[post]](https://openai.com/index/introducing-swe-bench-verified/)
post
364
[[Vendrow+ 2025]](https://arxiv.org/abs/2502.03461)
Do Large Language Model Benchmarks Test Reliability?
Joshua Vendrow, Edward Vendrow, Sara Beery, Aleksander Madry
2025-02-05
When deploying large language models (LLMs), it is important to ensure that these models are not only capable, but also reliable. Many benchmarks have been created to track LLMs' growing capabilities, however there has been no similar focus on measuring their reliability. To understand the potential ramifications of this gap, we investigate how well current benchmarks quantify model reliability. We find that pervasive label errors can compromise these evaluations, obscuring lingering model failures and hiding unreliable behavior. Motivated by this gap in the evaluation of reliability, we then propose the concept of so-called platinum benchmarks, i.e., benchmarks carefully curated to minimize label errors and ambiguity. As a first attempt at constructing such benchmarks, we revise examples from fifteen existing popular benchmarks. We evaluate a wide range of models on these platinum benchmarks and find that, indeed, frontier LLMs still exhibit failures on simple tasks such as elementary-level math word problems. Analyzing these failures further reveals previously unidentified patterns of problems on which frontier models consistently struggle. We provide code at https://github.com/MadryLab/platinum-benchmarks
365
![](https://cs336.stanford.edu/lectures/var/files/image-a1149b095a48ea306dcdea342363230c-https_pbs_twimg_com_media_GjICXQlWkAAYnDS_format_jpg_name_4096x4096)
366
![](https://cs336.stanford.edu/lectures/var/files/image-306ae3f862cee7c7842c0b29af3a2f5c-https_pbs_twimg_com_media_GjICcGQXYAAM4o1_format_jpg_name_4096x4096)
367
[[Zhu+ 2025]](https://arxiv.org/abs/2507.02825)
Establishing Best Practices for Building Rigorous Agentic Benchmarks
Yuxuan Zhu, Tengjun Jin, Yada Pruksachatkun, Andy Zhang, Shu Liu ... (15 more) ... Sarah Schwettmann, Matei Zaharia, Ion Stoica, Percy Liang, Daniel Kang
2025-07-03
Benchmarks are essential for quantitatively tracking progress in AI. As AI agents become increasingly capable, researchers and practitioners have introduced agentic benchmarks to evaluate agents on complex, real-world tasks. These benchmarks typically measure agent capabilities by evaluating task outcomes via specific reward designs. However, we show that many agentic benchmarks have issues in task setup or reward design. For example, SWE-bench Verified uses insufficient test cases, while TAU-bench counts empty responses as successful. Such issues can lead to under- or overestimation of agents' performance by up to 100% in relative terms. To make agentic evaluation rigorous, we introduce the Agentic Benchmark Checklist (ABC), a set of guidelines that we synthesized from our benchmark-building experience, a survey of best practices, and previously reported issues. When applied to CVE-Bench, a benchmark with a particularly complex evaluation design, ABC reduces the performance overestimation by 33%.
368
[[post]](https://transluce.org/introducing-docent)
post
369
370
371def how_to_think_about_evaluation():
372
373
374
375
376
377
378
379
380
381
382
383
384
385
![](https://cs336.stanford.edu/lectures/images/karpathy-nanogpt-speedrun.png)
[[post]](https://x.com/karpathy/status/1846790537262571739)
post
386
387
388
389
390
391
392
393if __name__ == "__main__":
394 main()
