
Testing statistical hypotheses
Introduction
Nova: Picture this: it is 1948 at the University of California, Berkeley. A young professor named Erich Lehmann is teaching a graduate course on hypothesis testing. In the audience sits a student named Colin Blyth, furiously taking notes, then going home each night to rewrite them into polished prose. Those notes, spanning just 163 pages in a flimsy paper cover, would become the seed of what is now one of the most cited books in all of statistics, with over sixteen thousand citations and counting. From its mimeographed underground origins to a thousand-page, two-volume fourth edition, E. L. Lehmann's Testing Statistical Hypotheses is nothing short of a legend. Today, we are diving into the story of this monumental book, the man behind it, and why it still matters.
Nova: Fair question! And I promise by the end of this conversation, you will see why this book is different. Think of it this way: before Lehmann's book, the world of hypothesis testing was a bit like a brilliant but disorganized workshop. You had geniuses like Fisher, Neyman, and Pearson inventing powerful tools, but there was no coherent manual. Lehmann's book provided the manual. It took all those scattered insights and wove them into a unified theoretical framework. One reviewer said Lehmann had a talent for clearing the fog and building a coherent structure. That is exactly what Testing Statistical Hypotheses did.
Nova: It is a wonderful origin story. Lehmann himself wrote a charming essay in 1997 called "Testing Statistical Hypotheses: The Story of a Book." He explains that Colin Blyth wrote up those meticulous notes, Lehmann reviewed them, they got mimeographed and sold at cost through the Berkeley Statistical Laboratory. As students graduated and took jobs at other universities, they started using those notes in their own courses. Orders began arriving from other colleges. It became what Lehmann called an "underground text."
Nova: Exactly! And it took ten more years, until 1959, for Lehmann to expand those notes into a proper book with Wiley. When asked what took so long, he summed it up in one word: rewriting. He said there was no page, section, or chapter that he could not improve a few months later, and improve again a few months after that. His advisor, the legendary Jerzy Neyman, finally had to tell him, "You have dawdled long enough. Get off the dime!"
E. L. Lehmann's Journey
The Man Behind the Masterpiece
Nova: Erich Leo Lehmann was born in Strasbourg in 1917 to a Jewish family. He grew up in Frankfurt, Germany, until 1933, when the Nazis came to power. His family fled to Switzerland. Imagine being sixteen years old and having your entire life uprooted. He finished high school in Zurich, studied mathematics at Trinity College, Cambridge, and then emigrated to the United States, arriving in New York in late 1940 with essentially no prior degree to his name.
Nova: Yes. He enrolled at Berkeley in 1941 and earned his MA in mathematics in 1942 and his PhD in 1946, studying under Jerzy Neyman, who co-created the Neyman-Pearson framework. So Lehmann learned hypothesis testing literally at the source. He also served as an operations analyst for the U. S. Air Force on Guam during World War II. And after the war, he became part of an extraordinary intellectual community at Berkeley that included Charles Stein, Joseph Hodges, and Henry Scheffé.
Nova: Precisely. Lehmann himself said he wove together two threads. The first was the Neyman-Pearson theory, with its core concepts of significance level, power, similarity, unbiasedness, and optimality. The second was Abraham Wald's decision theory, including ideas like minimaxity and admissibility, which he learned from Charles Stein. These two threads, he wrote, "easily combined into an integrated whole." That integration became the backbone of the book.
Nova: Absolutely. He is one of the eponyms of the Lehmann-Scheffé theorem on completeness of sufficient statistics, and of the Hodges-Lehmann estimator. He was a towering figure in nonparametric hypothesis testing. He won three Guggenheim Fellowships, was elected to the National Academy of Sciences, received honorary doctorates from Leiden and Chicago, and supervised over forty doctoral students, many of whom became leaders in the next generation. Peter Bickel, Frank Hampel, Allan Birnbaum — these are household names in statistics, and they all studied under Lehmann.
Nova: He had a great sense of humor. He also had a serious literary side. Lehmann wrote poetry, translated short stories by German authors like Adalbert Stifter into English, and published a professional autobiography called "Reminiscences of a Statistician: The Company I Kept." Later in life, he even wrote a history book about Fisher and Neyman. He was not just a mathematician. He was a humanist.
What the Book Actually Covers
The Architecture of a Classic
Nova: Let us get into the substance. What is actually inside Testing Statistical Hypotheses? At its core, the book is about optimality — finding the best possible test for a given statistical question. This traces back to the Neyman-Pearson lemma of 1933, which Lehmann called "mathematically quite elementary but with crucial statistical consequences."
Nova: Imagine you have a simple hypothesis, like "this coin is fair," versus a simple alternative, like "this coin lands heads sixty percent of the time." You want a test that controls your false positive rate — say, no more than five percent — while maximizing your chance of detecting the alternative when it is true. Neyman and Pearson proved that the likelihood ratio test is the optimal solution to that problem. It is the most powerful test you can possibly construct.
Nova: Exactly. And that is where Lehmann's book really shines. He systematically walks through what happens when you face complications. What if you want a test that works against many alternatives, not just one? In very rare cases, you can find a uniformly most powerful test, a UMP test, that beats everything else against all alternatives simultaneously. But those cases are rare indeed.
Nova: This is where the book's deep architecture comes in. Lehmann presents a series of strategies. One is unbiasedness — requiring that your test's power never drops below its significance level. Neyman and Pearson showed that the two-sided t-test is UMP among all unbiased tests. Another strategy is invariance: if your problem has a symmetry, you restrict attention to tests that respect that symmetry. This leads to the Hunt-Stein theorem and accounts for many standard tests like the F-test in analysis of variance.
Nova: That is a great way to put it. The book is organized into two major parts. Part one covers small-sample theory — what you can prove holds exactly, for any sample size. Part two covers large-sample theory, where you rely on asymptotic approximations as your sample size grows. The early editions focused more on small-sample theory, but with each revision, the large-sample treatment grew more sophisticated.
Nova: The first edition in 1959 was 369 pages, published by Wiley. The second in 1986 grew substantially. Then came the third edition in 2005, co-authored with Joseph P. Romano of Stanford, which added extensive material on asymptotic optimality and multiple testing. The fourth edition in 2022 is so large it had to be split into two volumes totaling over a thousand pages, with around nine hundred problems. Volume one covers finite-sample theory, volume two covers large-sample theory, and it includes entirely new chapters on multiple hypothesis testing, high-dimensional testing, permutation and randomization tests, and testing moment inequalities.
Nova: And that is exactly why it has endured. Each edition incorporated the most important developments in hypothesis testing since the previous one. The fourth edition, for example, reflects the explosion of interest in multiple testing due to genomics and big data, where you are testing thousands or millions of hypotheses simultaneously and need to control things like the false discovery rate.
Rigor, Legitimacy, and a Surprising Russian Story
The Measure Theory Gamble
Nova: Here is one of my favorite anecdotes from Lehmann's account of the book. When he was writing it, he faced a dilemma: what mathematical level should he pitch it at? The natural framework for a rigorous treatment of hypothesis testing is measure theory. But measure theory is notoriously abstract and difficult. He worried it would scare away readers.
Nova: Exactly. But Lehmann made a fascinating decision. He included a brief introduction to measure theory at the beginning of the book, giving the principal definitions and results, and then used them where needed in the rest of the text. He figured that even if readers skipped the measure theory, the statistical ideas would still come through clearly.
Nova: It led to a completely unexpected consequence. When a Russian translation of the book appeared years later, Lehmann heard from colleagues that his inclusion of measure theory actually helped legitimize statistics in the eyes of Russian mathematicians. They thought, well, if this statistics book is based on measure theory, there must be something to it. In the Soviet mathematical hierarchy, rigor grounded in measure theory gave statistics a credibility it had previously lacked.
Nova: And it speaks to Lehmann's care as a writer. He was not just dumping theorems on the page. He was deeply thoughtful about pedagogy, about who his readers would be and what they needed. He famously wrote and rewrote endlessly. He even talked about the agony of discovering errors after publication. Days after the first edition appeared, he found a serious mistake in one of the figures — the very figure the publisher had chosen for the dust jacket.
Nova: He described letters from readers that would begin, "I have difficulty following your proof..." which he learned to interpret as the polite academic way of saying, "You botched it again." But he took all of it with good humor and kept improving the book.
Nova: This is a crucial part of the story. R. A. Fisher and Jerzy Neyman had a bitter philosophical disagreement about the foundations of hypothesis testing. Fisher favored significance testing with p-values and did not accept the idea of alternative hypotheses or power calculations. Neyman and Pearson developed the formal framework of Type I and Type II errors, rejecting or accepting hypotheses. The two camps were often at odds. Lehmann, as Neyman's student, was firmly in the Neyman-Pearson tradition, but he was also deeply respectful of Fisher. In 1993, he published a famous paper called "The Fisher, Neyman-Pearson Theories of Testing Hypotheses: One Theory or Two?" where he argued that the two approaches could be seen as complementary rather than contradictory.
From Genomics to A/B Testing
Why This Book Still Matters
Nova: Let us talk about why Testing Statistical Hypotheses matters for today's world. When Lehmann first wrote the book, the primary applications were in agriculture, industrial quality control, and the social sciences. But hypothesis testing has exploded into virtually every domain of modern life.
Nova: Take genomics. When scientists sequence a human genome, they are comparing gene expression levels across thousands of genes between, say, cancer patients and healthy controls. That means running thousands of hypothesis tests simultaneously. If you use a standard five-percent significance level for each test, you will drown in false positives. The multiple testing framework developed in the later editions of Lehmann and Romano's book provides the rigorous foundation for controlling the false discovery rate. It is literally the mathematical backbone behind modern drug discovery.
Nova: Absolutely. Or consider A/B testing. Every time a tech company tests two versions of a webpage to see which drives more clicks, they are performing a hypothesis test. The asymptotic theory covered in the book justifies using large-sample approximations to compute p-values and confidence intervals in these settings. The randomization and permutation tests covered in the fourth edition are increasingly important in analyzing experiments where traditional parametric assumptions do not hold.
Nova: That is a really important point. Lehmann was aware of these concerns even decades ago. In his essay on the history of optimality, he discussed John Tukey's criticism that the Neyman-Pearson framework could become a mechanical ritual rather than a thoughtful scientific tool. Lehmann's response, implicit in his entire approach, was that understanding the theoretical foundations deeply is the best antidote to misuse. If you understand why a test is optimal, what assumptions it requires, and what its limitations are, you are far less likely to apply it mindlessly.
Nova: I would go further. The book also covers what happens when standard assumptions break down. The chapters on nonparametric and permutation-based methods, on robustness, and on testing when parameters are not identified all speak directly to the challenges that real data analysis faces. This is not a cookbook. It is a deep exploration of the logic of statistical inference.
Nova: Exactly. Joseph Romano, Lehmann's co-author for the third and fourth editions, has been instrumental in keeping the book at the frontier. Romano himself is a major figure who invented statistical methods like subsampling and the stationary bootstrap, and has contributed to multiple hypothesis testing methodology. Under his stewardship, the book has remained not just a reference work but a living, growing document.
Lehmann's Philosophy and Legacy
The Human Side of a Monumental Work
Nova: I want to circle back to something that makes this story so compelling: the human dimension. Lehmann was not just a theorem machine. He was a refugee who found a new home in American academia, a mentor who shaped generations of statisticians, and a person who genuinely loved language and literature.
Nova: Very few. His colleague Peter Bickel wrote that Lehmann was "kind and generous of spirit, had an unusual sensitivity to the feelings of others and a great astuteness about the world." Bickel also noted that while Lehmann achieved every major honor in the field, he "served reluctantly but very effectively as Department Chair." He preferred writing and mentoring to administration.
Nova: Yes. In 1977 he married Juliet Popper Shaffer, a distinguished statistician in her own right, who specialized in multiple comparison procedures. She had come to Berkeley as a visiting scholar, and they ended up spending the rest of his life together. She was also active in the later stages of some of his work.
Nova: I think it is this: that hypothesis testing is not a bag of tricks or a collection of ad hoc procedures. It is a coherent, principled framework where you can rigorously justify why one test is better than another. The central organizing question is always: among all tests that control the probability of false rejection at a given level, which one maximizes the probability of detecting a true effect? That question, asked systematically across different statistical settings, is the unifying thread of the entire book.
Nova: That is beautifully put. And that, I think, is why the book has had such a long life. It does not just teach you how to test hypotheses. It teaches you how to think about testing hypotheses.
Nova: Honestly, it is a graduate-level text that assumes substantial mathematical preparation. It is not light reading. But I would say this: even if you never open the book, knowing that it exists, knowing that there is a rigorous foundation underneath all the hypothesis tests you encounter in news articles, medical studies, and business analytics, should give you a certain confidence. And if you are a data scientist, an economist, a biologist, or anyone whose work involves drawing conclusions from data, there is enormous value in at least knowing what this book represents: the gold standard for thinking carefully about statistical evidence.
Conclusion
Nova: We have covered a lot of ground. From a set of student notes mimeographed in 1949 to a thousand-page, two-volume fourth edition in 2022, Testing Statistical Hypotheses has been the definitive guide to the theory of hypothesis testing for over six decades. It was written by a man who fled the Nazis, studied under the founders of modern statistics, and devoted his life to bringing clarity and rigor to a field that was once chaotic and fragmented.
Nova: Erich Lehmann died in 2009 at the age of 91, but his book lives on, now shepherded by Joseph Romano. With over sixteen thousand citations, it remains one of the most referenced works in all of statistics. More importantly, it embodies an ideal: that deep theoretical understanding is not opposed to practical application — it is the foundation that makes application trustworthy.
Nova: It is a reminder that great things can have humble beginnings, and that the pursuit of clarity, rigor, and understanding is one of the most generous gifts a scholar can give to the future.