Bioinformatics for Beginners
Genes, Genomes, Molecular Evolution, Databases and Analytical Tools
Introduction
Nova: Picture this. You're a first year biology student, your lecturer hands out a reading list, and near the top it says 'Bioinformatics for Beginners, by T. K. Attwood.' So you search for it, and the internet gives you... a shrug.
Nova: That's the funny thing. The title is crystal clear, but the author is a bit of a ghost. It turns out there are two different books hiding inside that one citation, and neither is exactly what the syllabus claims.
Nova: Welcome to the case of the vanishing textbook. Today we're untangling one of the most quietly widespread mix-ups in biology education. And while we're at it, we'll get a proper tour of bioinformatics itself: what it is, who built it, and why the books we hand to beginners matter.
Nova: Exactly. It starts with a name, T. K. Attwood. She's a real person, a pioneering bioinformatician at the University of Manchester. But the book usually credited to her as 'Bioinformatics for Beginners' may not be the one she actually wrote.
Key Insight 1: Untangling the Title
The Case of the Two Books
Nova: Let's start with the paperwork. There is a real book called 'Bioinformatics for Beginners.' It was published in 2014 by Academic Press, an imprint of Elsevier, and its full title is 'Bioinformatics for Beginners: Genes, Genomes, Molecular Evolution, Databases and Analytical Tools.'
Nova: The problem is the name on the cover. It's not T. K. Attwood. The author is Supratim Choudhuri, a toxicologist at the U. S. Food and Drug Administration's Center for Food Safety and Applied Nutrition.
Nova: It actually makes a lot of sense once you meet him. Choudhuri has published for decades in molecular toxicology, genomics and epigenetics. He wrote the book from the perspective of a working scientist who needs to use databases and tools daily, not from the perspective of a software engineer.
Nova: She fits in because she wrote a different, much older and very famous book. In 1999, Teresa Attwood and David Parry-Smith published 'Introduction to Bioinformatics' with Longman, as part of the 'Cell and Molecular Biology in Action' series. It was a final year undergraduate text focused on two areas: genomics and protein sequence analysis.
Nova: That's exactly the mess. University reading lists all over the world list 'Bioinformatics for Beginners, T. K. Attwood' as a reference. Someone remembered the friendly beginner title and the famous bioinformatics author, and fused them together.
Nova: In practice they're probably fine, but the mix-up is worth noticing because it tells us something about how knowledge travels. The title 'for beginners' is what students remember. The name Attwood is what lecturers remember. And the actual 2014 book, Choudhuri's, sometimes gets lost entirely.
Nova: Precisely. And here's a stat to anchor it: Attwood's 'Introduction to Bioinformatics' has been cited around 396 times according to Manchester's research records, and it went through multiple printings. Choudhuri's 'Bioinformatics for Beginners' is newer and widely used in graduate courses. Both are legitimate, and both are completely different animals.
Nova: Let's do it.
Key Insight 2: Teresa K. Attwood
The Scientist Behind the Name
Nova: Teresa K. Attwood goes by Terri, and she's been at this a very long time. She did her bachelor's in biophysics at the University of Leeds in 1982, then a PhD in biophysics two years later, in 1984.
Nova: Here's the surprise: chromonic mesophases. That's a kind of liquid crystal. Her early work looked at how certain anti-asthmatic drugs form these lyotropic liquid crystal phases. So she started out studying the physics of molecules stacking themselves.
Nova: It is, and yet it makes sense. Liquid crystals are all about how molecules arrange themselves into patterns. A decade later, she'd be building pattern-finding tools for proteins. The obsession with structure and pattern never left.
Nova: Inspired by the PROSITE database of protein signatures, Attwood developed a method called protein fingerprinting. In 1994, she and Mark Beck published the PRINTS database, which describes proteins by groups of conserved motifs rather than a single one. That fine-grained approach is her signature contribution.
Nova: Great analogy. And that idea snowballed. In 1997, she teamed up with Amos Bairoch and others to unify protein family classification, which grew into InterPro, the massive integrated resource for protein families, domains and sites. Today InterPro is one of the most used resources in the field, run jointly with the European Bioinformatics Institute.
Nova: Right. She also co-developed tools like CINEMA, an early interactive editor for multiple sequence alignments, and UTOPIA, a set of tools for visualising and linking biological data with the scientific literature. She even worked on the Semantic Biochemical Journal, where you can hover over a protein name in a paper and see its data live.
Nova: And the education side is just as deep. She led EMBER, a European multimedia bioinformatics education project, and in 2012 she helped found GOBLET, the Global Organisation for Bioinformatics Learning, Education and Training. She's literally been organising how the world teaches this subject.
Nova: Exactly. And here's the kicker: she wrote three textbooks, not one. 'Introduction to Bioinformatics' in 1999, then 'Bioinformatics and Molecular Evolution' with Paul Higgs in 2005, and 'Bioinformatics Challenges at the Interface of Biology and Computer Science: Mind the Gap' in 2016.
Nova: Nicely put. Now let's actually crack open the books and see what they teach.
Key Insight 3: Two Different Entry Ramps
What's Actually Inside
Nova: Attwood and Parry-Smith's 'Introduction to Bioinformatics' is a time capsule from 1999, and that's not a criticism. It opens with the internet and the world wide web, because back then, just getting a biologist online was half the battle.
Nova: Exactly. The book walks students through primary, composite and secondary databases, then dives into genomics and protein sequence analysis. It's lean, around 240 pages, and it launched at about twenty pounds.
Nova: Right. It assumes you're a final year undergraduate or starting a master's, and it holds your hand through the concepts. Now contrast that with Choudhuri's 2014 book. It has nine chapters and over a hundred figures, and it's built around a different spine: molecular evolution.
Nova: Because Choudhuri argues that you can't understand why sequences look similar until you understand descent, mutation, recombination and the neutral theory. So chapters one and two are fundamentals of genes and genomes, then fundamentals of molecular evolution. He lays the biological context before touching a single tool.
Nova: And that's exactly the philosophy of the book. Chapter three covers sequencing technologies, from Sanger to pyrosequencing to next generation platforms. Chapter four is the history of the field itself, starting with Margaret Dayhoff, Richard Eck and Robert Ledley in the 1960s.
Nova: She's essential. Dayhoff built the first protein sequence database in the mid 1960s and created the point accepted mutation matrices still used in sequence alignment today. Choudhuri treats her as the starting point of the discipline.
Nova: Then chapters five through nine are the hands-on core: databases and data formats, BLAST and FASTA for similarity searching, nucleic acid analyses, protein analyses, and finally phylogenetic tree building.
Nova: Yes, and Choudhuri is careful to explain that both are fast, heuristic versions of the rigorous Smith-Waterman algorithm. He spends real time on sequence identity versus similarity versus homology, a distinction beginners constantly get wrong.
Nova: Exactly, and a good beginner book fixes that early. The Attwood book, by the way, does something similar on the protein side, teaching you to read motifs and fingerprints rather than just blast and hope.
Nova: That's the cleanest summary. One is a focused classic, the other is a comprehensive modern tour. Both end up at the same place: teaching you to ask biological questions with computers.
Nova: Let's get into that.
Key Insight 4: The Education Gap
Why Beginners Need a Guide
Nova: Here's a number to chew on. The Human Genome Project finished in 2003, but its working draft landed in 2001. Attwood wrote her textbook two years before that draft, when databases were small enough to browse by hand.
Nova: That's the core tension. The tools are now free and powerful, but they're overwhelming. Attwood has spent her career warning about what she calls knowledge lost in the literature and data landslide.
Nova: She wrote a paper with that exact theme in 2009, arguing that we have incredible data but we're losing the ability to make sense of it. That's why she kept pushing education, founding GOBLET and building training networks across the world.
Nova: Choudhuri's book approaches the same problem from a different angle. He wrote it for the working biologist who is not a programmer, someone who needs to do real analysis with free web tools but doesn't want to be buried in theory.
Nova: Exactly. His book is unusual because a practicing toxicologist wrote it, someone who actually sits at the interface of genomics and regulation. He even brought in a colleague, Michael Kotewicz, to contribute the section on optical mapping of DNA.
Nova: Right. And that's a theme across both authors: meet the beginner where they are. Attwood's book literally starts by teaching you what the web is. Choudhuri's book starts by teaching you why DNA mutates and evolves. Neither assumes you already speak the jargon.
Nova: Treating bioinformatics as button clicking. Searching a database without understanding the biology behind the similarity, or trusting a tool's output without knowing its limits. Both books push back on that by teaching principles first, tools second.
Nova: And that's the real reason these books still matter, even in an age of YouTube and free MOOCs. A well-structured book forces you to build the mental model, not just copy steps.
Nova: That's the perfect place to land our conclusion.
Conclusion
Nova: Let's pull the threads together. First, the mystery: there is no single book called 'Bioinformatics for Beginners' by T. K. Attwood. The friendly title belongs to Supratim Choudhuri's 2014 Elsevier book, while Attwood's classic is 'Introduction to Bioinformatics,' written with David Parry-Smith in 1999.
Nova: Second, the person. Teresa Attwood went from liquid crystal physics to building PRINTS and InterPro, and she spent decades making sure the field teaches itself properly through GOBLET and EMBER. She's a founder, not just an author.
Nova: Third, the lesson. The best beginner resources, both of these included, teach you principles before tools. They teach you why sequences are similar before they teach you to click BLAST. They treat homology, evolution and databases as the real curriculum.
Nova: If you want the modern, hands-on, full-spectrum tour, go with Choudhuri's 'Bioinformatics for Beginners.' If you want the classic, protein-focused introduction written by a field pioneer, grab Attwood and Parry-Smith's 'Introduction to Bioinformatics.' Honestly, reading both would be a fantastic start.
Nova: Bioinformatics is not about the software. It's about asking biological questions that the software can answer. Get the biology and the evolution right, and the buttons will take care of themselves.
Nova: And a reminder that even in a digital age, a well-chosen book is still one of the best on-ramps into a field.
Nova: This is Aibrary. Congratulations on your growth!