Podcast thumbnail

Introduction to bioinformatics

11 min
4.7

Introduction

Nova: Welcome to Aibrary, the podcast where we crack open the most influential books in science and transform them into conversations you actually want to listen to. I'm Nova, your knowledgeable guide.

Nova: That's right. We're diving into Introduction to Bioinformatics by Arthur M. Lesk, now in its fifth edition, published by Oxford University Press. This book has been the go-to introductory text for bioinformatics for over two decades. And here's a surprising fact to start us off: its author, Arthur Lesk, isn't just a textbook writer. He co-authored one of the most cited papers in structural biology, a 1986 EMBO Journal paper with Cyrus Chothia that has been cited over 3,600 times. It fundamentally changed how scientists predict protein structures.

Nova: Exactly. Lesk brings genuine research authority to every page. He was at the MRC Laboratory of Molecular Biology in Cambridge, he worked at the European Molecular Biology Laboratory in Heidelberg, and he's now a professor at Penn State. He wrote the first computer program to generate schematic diagrams of protein structures. In 2023, he received the Carl Brändén Award for outstanding contributions to protein science and education.

Nova: That's exactly what we're going to unpack today. From genomes to AI, from evolutionary trees to drug discovery, Lesk's Introduction to Bioinformatics is a masterclass in making the complex accessible. Let's get into it.

The Genesis and Evolution of Lesk's Classic

How One Book Defined a Field

Nova: Let's start with the big picture. When Lesk first published Introduction to Bioinformatics in 2002, the Human Genome Project had just delivered its draft sequence. Biology was drowning in data and desperately needed people who could swim. This book arrived at exactly the right moment.

Nova: You're right to ask. The key difference was accessibility. Lesk explicitly wrote for a biological audience with no prior programming knowledge. He doesn't throw proofs at you. He doesn't assume you know dynamic programming or graph theory. One reviewer on Goodreads put it perfectly: the book contains virtually no mathematical sophistication, and complex topics like Hidden Markov Models are explained in an unintimidating, intuitive manner.

Nova: They aren't, but Lesk makes them approachable. And the book has evolved dramatically. The first edition was about 320 pages. The fifth edition, published in 2019, runs 432 pages and includes entirely new chapters on artificial intelligence and machine learning, plus new content on next-generation sequencing, epigenomics, gene editing, and single nucleotide variants.

Nova: That's a perfect bioinformatics analogy. And the mutations have been adaptive. The fifth edition's table of contents tells the story: it moves from genetics to genomes, to the panorama of life, into alignments and phylogenetic trees, structural bioinformatics, scientific publishing, AI and machine learning, systems biology, metabolic pathways, and control mechanisms.

Nova: And what reviewers consistently praise is that Lesk doesn't just tell you what the tools are. He tells you why you're learning them and what you can do with them. One reviewer wrote that the book is organized around tools of the trade rather than grandiose theory, which makes it immediately useful for undergraduates and researchers new to the field.

Nova: Great question. The same Goodreads reviewer, who came from a computer science background, said biological concepts were sufficiently explained and that he found the book a cinch to read. So it works in both directions, though it's definitely tilted toward the biology audience. Lesk himself has an interesting background: Harvard undergraduate, Princeton PhD, Cambridge master's, and he's worked across chemistry, molecular biology, and computation. He embodies the interdisciplinary nature of bioinformatics.

The Core Toolkit: Alignment, BLAST, and Phylogenetics

From Sequence to Significance

Nova: Let's dig into the real meat of the book. Lesk states clearly that sequence alignment is, and I'm quoting, THE basic tool of bioinformatics. Chapter four of the fifth edition is dedicated to alignments and phylogenetic trees, and it's where most students spend the bulk of their mental energy.

Nova: Imagine you have two strings of DNA letters, maybe thousands of characters long. Sequence alignment is the process of arranging those sequences to identify regions of similarity. These similarities might indicate functional, structural, or evolutionary relationships. It sounds simple, but it's computationally intense. Lesk walks you through dot plots, pairwise alignment, multiple sequence alignment, profiling, and then the superstar of the field: BLAST.

Nova: Exactly. And Lesk devotes significant attention to BLAST and its variants, including PSI-BLAST, which stands for Position-Specific Iterative BLAST. The beauty of Lesk's approach is that he situates each tool, explaining its advantages and limitations. He wants you to understand not just how to run BLAST, but when to use it and when not to trust the results.

Nova: This is where Lesk's research expertise really shines. His 1986 paper with Chothia showed that the relationship between sequence similarity and structural similarity is nonlinear. Two proteins can have very similar structures even when their sequences have diverged significantly. So a BLAST search might miss a structural relative. Lesk is careful to emphasize the difficulty of inferring homology from sequence similarity alone. He warns against making assumptions about mutation rates. These caveats come from someone who has spent decades working precisely on these problems.

Nova: That's beautifully put. And there's a wonderful practical tip that one reviewer highlighted. Lesk writes that visual examination of multiple sequence alignment tables is one of the most profitable activities a molecular biologist can undertake away from the lab bench. He even tells you: don't even think about not displaying them with different colors for amino acids of different physiochemical types.

Nova: Exactly. The phylogenetic trees chapter then extends from alignment into evolutionary reconstruction. You learn how to build trees, how to interpret them, and crucially, how not to overinterpret them. The chapter connects molecular data to the grand narrative of evolution, which Lesk frames beautifully in his chapter on the panorama of life.

Structural Bioinformatics, Drug Discovery, and Lesk's Own Legacy

The Shape of Life Itself

Nova: Chapter five of the book covers structural bioinformatics and drug discovery, and this is where Lesk's personal research story becomes inseparable from the textbook narrative.

Nova: Yes, and it's fascinating. Lesk and Chothia discovered something fundamental: the relationship between changes in amino acid sequence and changes in protein structure. They found that when two protein sequences are more than about 50 percent identical, their three-dimensional structures are almost always very similar. Below that threshold, structural divergence accelerates in a nonlinear way.

Nova: Precisely. And this discovery provided the quantitative foundation for homology modeling, which remains one of the most successful and widely used methods of protein structure prediction. It's the basis for tools like MODELLER and SWISS-MODEL that researchers use every day.

Nova: Lesk covers this directly. If you know a protein's structure and you know it's involved in a disease, you can computationally screen millions of small molecules to find ones that might bind to it and alter its function. This is structure-based drug design. Lesk's own work on antibodies is a perfect case study. He and Chothia discovered the canonical-structure model for antibody binding sites, which supported the humanization of antibodies for cancer therapy.

Nova: Here's the story: researchers found that rats can raise antibodies against human cancers. But when you inject rat antibodies into human patients, the human immune system recognizes them as foreign and mounts a response, almost like an allergic reaction. The solution is to create hybrid molecules that are more human than rat, retaining the cancer-targeting ability while reducing the immune response. Lesk's structural work on antibody binding sites made it possible to predict which parts of the rat antibody could be swapped for human components without destroying the therapeutic activity.

Nova: It really is. And Lesk also developed one of the first computer programs to generate schematic diagrams of protein structures. Those beautiful ribbon diagrams you see in every molecular biology paper, with alpha helices as spirals and beta sheets as arrows, his program was among the pioneers of that visualization approach. He worked with Karl Hardman and built on Jane Richardson's classification scheme for ribbon diagrams.

AI, Systems Biology, and the Future of Bioinformatics

Machines That Learn Biology

Nova: The fifth edition added something that would have seemed almost science fiction when the first edition came out: an entire chapter on artificial intelligence and machine learning in bioinformatics.

Nova: It's one of the things that makes the fifth edition so valuable. Lesk integrates AI and machine learning not as some futuristic add-on, but as a natural extension of the computational approaches he's been describing all along. The chapter covers how machine learning methods are applied to function prediction, to analyzing genomic data, to classifying sequences, and to the kind of pattern recognition problems that pervade bioinformatics.

Nova: Absolutely. And what's interesting is how Lesk positions it. He doesn't present AI as replacing traditional bioinformatics. He shows it as the next logical step. Chapter eight then moves into systems biology, which looks at biological systems holistically rather than one gene or one protein at a time. Chapter nine covers metabolic pathways, and chapter ten addresses the control of organization in biological systems.

Nova: It really is. And the fifth edition also includes substantial new material on next-generation sequencing, epigenomics, and the bioinformatics of gene editing. The CRISPR revolution has created entirely new computational challenges: how do you predict off-target effects? How do you design guide RNAs? How do you analyze the results of a gene editing experiment? Lesk covers all of this.

Nova: It is unusual, and it's one of the things that makes Lesk's book distinctive. He recognizes that a huge part of bioinformatics is knowing how to access and navigate the scientific literature. He covers databases like PubMed, the structure of scientific articles, how to search effectively, and how to extract information from the literature using text mining and natural language processing. In the era of information overload, this chapter is arguably more important than ever.

Nova: Exactly. And throughout the book, Lesk includes frequent examples, self-test questions, problems, and exercises. The online resources include something called Weblems, web-related problems that get students actually using the tools they're reading about. It's active learning embedded in a textbook format.

Conclusion

Nova: So, Stella, after exploring Lesk's Introduction to Bioinformatics across its five editions and ten chapters, what stands out to you?

Nova: That's the magic of it, isn't it? The book is praised for its lucidity. One reviewer said you can actually read it from cover to cover and remember what you read. In an era of dense, impenetrable textbooks, that's a genuine achievement.

Nova: If there's one takeaway for our listeners, it's this: bioinformatics can seem intimidating, with its algorithms and databases and command lines. But Lesk's book proves that with the right guide, anyone can enter this world. And the guide happens to be one of the scientists who built it.

Nova: Lesk opens his preface with a quote from Shakespeare's Antony and Cleopatra: In nature's infinite book of secrecy, a little I can read. That humility, combined with deep expertise, is what makes this book special. Lesk shows you not just what we can read in nature's book, but how we read it, and what tools make that reading possible.

00:00/00:00