Next-Generation Sequencing (NGS): Principle, Steps, and the Illumina Workflow
Next-generation sequencing reads millions of DNA fragments at once. Learn the principle, the four-step Illumina workflow (library prep, cluster generation, sequencing by synthesis, data analysis), and how base quality (Q30) is scored.
On this page
The defining feature of next-generation sequencing is scale: instead of reading one DNA fragment at a time as older methods do, it reads millions to billions of fragments at the same time. This is why it is also called massively parallel sequencing, and it is what reduced the time to sequence a human genome from years to about a day.
Next Generation Sequencing (NGS) technology has transformed how clinical researchers and scientists think about genetics, as it assesses multiple genes in a single assay. It can sequence an entire or particular genome of interest within a short period. Several different NGS platforms use other sequencing technologies. Still, one common thing is ‘all NGS platforms execute sequencing of millions of small fragments of DNA in parallel.’ Bioinformatics analyses are then used to piece together these fragments.
Principle of Next-Generation Sequencing
Next Generation Sequencing technology is similar to Capillary Electrophoresis(CE) sequencing, where DNA polymerase incorporates fluorescently labeled, reversibly terminated nucleotides into a growing strand during sequential cycles of DNA synthesis, one base at a time. During each process, the nucleotides are identified by fluorophore excitation at the addition of each nucleotide. The major difference is that instead of sequencing a single DNA fragment, NGS simultaneously extends this process across millions of fragments. NGS delivers high accuracy, an elevated yield of error-free readings, and a high percentage of base calls above Q30.
In short, the principle is massive parallelism: break the genome into millions of small fragments, copy each into a signal-strong cluster, and sequence all of them simultaneously, then use a computer to assemble the reads back into the full sequence.
Unlike Sanger sequencing, which reads one fragment at a time and stops each strand permanently, NGS reads millions of fragments at once and uses reversible terminators, so each strand is read one base at a time. The chain-termination chemistry that Sanger sequencing relies on is described in the separate article on Sanger sequencing.
Steps of Next-Generation Sequencing (Illumina Workflow)
Illumina sequencer includes four basic steps of sequencing: Library Preparation, Cluster Generation, Sequencing, and Analysis. After the isolation of DNA or cDNA (synthesized from RNA), they undergo all of these four basic steps.
At a glance, the four steps are:
- Library preparation: fragment the DNA and attach adapters (and sample barcodes).
- Cluster generation: fix fragments to a flow cell and copy each about a thousand times by bridge amplification, so each spot gives a strong signal.
- Sequencing by synthesis: read every cluster at once, adding one fluorescently labeled, reversibly blocked base per cycle and imaging after each base.
- Data analysis: convert the images to base calls, then align or assemble the millions of short reads into the final sequence.
Each step is described in detail below.
Library Preparation
The library preparation is vital in the sequencing process as it prepares the samples to be compatible with the sequencer. The samples, either DNA or cDNA, are fragmented by sonication or enzymatic restriction to obtain fragments of 200-500 bp in length. During fragmentation, each fragment gets an overhang A tail at the end, preparing them to ligate to the adapter sequence, which contains a ‘T’- base overhang complementary to the A-tail fragment.
Adapters are the sequence that contains the primer binding sites, index sequences, and the sequence that allows library fragments to attach to the flow cell lawn. These adapters are ligated in 5’ and 3’ ends. Alternatively, a process called tagmentation combines fragmentation and ligation reactions in a single step, increasing the efficiency of the library preparation process.
The primer binding sites in the adapters also allow specific enrichment of adapter-ligated fragments during the later PCR step.
Figure: The outer region in black and orange binds to the complementary sequence on the surface of the Illumina flow cell. Through which the individual library single strands are captured to be sequenced. The inner region of green and blue acts as a sequencing primer binding site which is used to read out the insert sequence during the actual sequencing process.Source:https://www.lexogen.com/rna-lexicon-next-generation-sequencing/
For each sample, unique adapters are used, and all the samples are pooled together in a single tube. After pooling together, it is loaded in the sequencer for further proceeding into sequencing.
Figure: Fragmenting DNA sample and ligating specialised adapters to both fragments ends to make NGS library. Source:https://www.cd-genomics.com/blog/principle-and-workflow-of-illumina-next-generation-sequencing/
Cluster Generation
The prepared library is loaded in a flow cell and placed in the sequencer. A flow cell is a glass slide with 1, 2, or 8 physically separated lanes coated with a lawn of surface-bound, adapter-complimentary oligos. Nowadays, mostly used flow cell is a patterned flow cell produced using semiconductor manufacturing technology. It has a glass substrate containing patterned nano wells with DNA probes that capture the prepared DNA strands for amplification during cluster generation.
Figure: Illumina Patterened flow cell. Source:https://www.illumina.com/science/technology/next-generation-sequencing/sequencing-technology/patterned-flow-cells.
Cluster generation begins once the samples are attached to the flow cell and bridge amplification occurs. During this process, the DNA fragments of the library get hybridized with one of the oligonucleotides present on the flow cell surface. This is then followed by generating a complementary strand by elongating the oligo attached to the flow cell by DNA polymerase. The original molecule is washed away, and the strand bends over like a bridge and attaches to the next oligo in the flow cell. This second oligo is complementary to another adapter sequence, and polymerase generates a complementary strand forming a double-stranded bridge. This bridge is denatured, resulting in two single-stranded copies of the molecule tethered to the flow cell. The process is repeated over and over and occurs simultaneously for millions of clusters resulting in clonal amplification of all.
After bridge amplification, the reverse strands are cleaved and washed off, leaving only the forward strands. The 3′ ends are blocked to prevent unwanted priming. When cluster generation is complete, the templates are ready for sequencing.
Figure: Bridge Amplification Process.
Sequencing by Synthesis
Sequencing occurs for every cluster on the flow cell at the same time. The components required for sequencing include the sequencing primer, DNA polymerase, and nucleotides that are each labeled with a fluorophore. However, based on the chemistry used in their respective machine (4-channel chemistry, 2-channel chemistry, and 1-channel chemistry), either all the nucleotides are labeled, or only a few nucleotides are labeled with the fluorophore
In 4-channel chemistry, all the nucleotides are labeled with four fluorescent dyes. Two-channel chemistry uses two different fluorescent dyes, while one-channel chemistry uses only one dye.
Figure: Four-, Two- and One- Channel Chemistry: Four Channel chemistry uses nucleotides labelled with four different dyes, Two channel chemistry uses two different fluorescent dyes and one channel Chemistry uses only one dye. Source: https://www.illumina.com/content/dam/illumina-marketing/documents/products/techspotlights/cmos-tech-note-770-2013-054.pdf
In the case of 4-channel chemistry, all the above components are passed through the flow cell, where sequencing primer anneals to its complementary location on the adapter, and DNA polymerase adds the complementary nucleotides labeled with a fluorophore.
Each labeled nucleotide also carries a reversible terminator, a blocking group that stops the polymerase after a single base is added. This pause allows the detector to record the fluorescence of that one base. The block and dye are then removed so the next base can be added in the following cycle. Because the block is reversible, the same strand is read one base at a time, cycle after cycle.
After adding each nucleotide, fluorophore in the cluster are excited by the light source and the characteristic fluorescent signal emitted are recorded. This process is called sequencing-by-synthesis.

How to read the nucleotide in NGS?
During sequencing, each cluster is read base by base to produce its sequence. Each cluster gives one read per sequencing direction: one read for single-read sequencing, or two (one from each end) for paired-end sequencing. To begin, the flow cell is flooded with the sequencing components (DNA polymerase, sequencing primer, and fluorescently labeled nucleotides). The sequencing primer anneals to the adapter, and the fragment is read from one end. This first read is called Read 1.
After Read 1, a short separate index read is usually performed to decode the sample barcode in the adapter, which is what allows pooled samples to be told apart later. For paired-end sequencing, the template is then regenerated on the flow cell and folds over to bind a surface oligo, and a second sequencing read (Read 2) is performed from the opposite end of the same fragment. Reading both ends improves how accurately the fragment can be aligned. In single-read sequencing, the fragment is read from one end only. Throughout, the fluorophore emitted after each base is detected for every cluster by the optical system to determine the base added.
Figure: Addition of fluorescently labelled nucleotide and identifying the fluorophore. Source: https://www.lexogen.com/rna-lexicon-next-generation-sequencing/
Data Analysis
After the sequencing is complete, the optical signals are translated to a nucleotide sequence called base calling. The accuracy of base calling is measured by the Phred quality score (Q score), the most common metric for assessing sequencing data quality. Q score indicates the probability that the given base is called incorrectly by the sequencer.
The Q score is logarithmically related to the base-calling error probability (P), by the formula Q = -10 log₁₀P. A higher Q score means a lower chance of error.
This Q score determines a good base or a bad one. The Quality Score and Base Calling Accuracy is shown below. The scoring Q30 is ideal for a range of sequencing applications.
Figure: Quality Score and Base Calling Accuracy. Source:https://www.illumina.com/documents/products/technotes/technote_Q-Scores.pdf
During the library preparation, each sample is given unique index sequences, which are called multiplexing; this allows large numbers of libraries to be pooled together and sequenced simultaneously in a single sequencing run. So, before data analysis, a process called demultiplexing occurs, which separates the sequences from pooled sample libraries based on their unique indexes. For each sample, reads with similar stretches of the bases are locally clustered. Forward and reverse reads are paired, creating contiguous sequences. These newly identified contiguous sequences are aligned to a reference sequence.
Figure: Multiplexing process is shown in A during Library Preparation, where unique indexes are provided to each sample. After library preparation, each sample is pooled together. The sequencing process is shown in C. After sequencing, the demultiplexing algorithm sorts the reads into different files according to their indexes. Source:https://www.illumina.com/content/dam/illumina-marketing/documents/products/illumina_sequencing_introduction.pdf
The number of reads covering each position of the reference genome is called the coverage depth. Higher coverage means each base has been read many times over, which makes the final call at that position more reliable. Comparing the aligned reads against the reference then reveals the differences and similarities, which are interpreted using bioinformatics tools.
Figure: Reads are aligned to the reference sequence. After the alignment, the differences between them can be identified. Source: Illumina sites
Following alignment, many analysis variations are possible, such as single nucleotide polymorphism (SNP) or insertion-deletion (indel) identification, read counting for RNA methods, phylogenetic or metagenomic analysis, and more.
Application of Next-Generation Sequencing
Next Generation Sequencing technology has a variety of applications. Some of these are mentioned below:
- Identifying pathogens directly from clinical samples. NGS can detect the organism causing an infection straight from a specimen such as blood, cerebrospinal fluid, or respiratory fluid, without waiting for culture. This is especially useful for organisms that grow slowly or do not grow in culture, for infections involving several organisms at once, and for cases where the patient has already received antibiotics. It has become valuable in difficult situations such as sepsis, meningitis, and infections in immunocompromised patients.
- Detecting antimicrobial resistance. By reading the genome of a pathogen, NGS can find the specific genes and mutations that make it resistant to particular antibiotics. This allows resistance to be predicted from the DNA itself, which can support faster and more targeted treatment decisions than waiting for conventional susceptibility testing alone.
- Tracking outbreaks and disease surveillance. Because whole-genome sequencing can tell apart strains that look identical by ordinary tests, it can show whether cases in an outbreak share the same source. Public health and hospital infection-control teams use this to trace food-borne outbreaks, follow the spread of hospital-acquired infections, and monitor resistant organisms across regions.
- Detecting and characterizing viruses. NGS is widely used to identify viruses, including newly emerging ones, and to follow how they change over time. It was central to identifying and tracking SARS-CoV-2 variants during the COVID-19 pandemic.
- Metagenomic sequencing. This approach sequences all the genetic material in a sample at once, without culturing or targeting specific organisms. It is used to study whole microbial communities, such as the gut or environmental microbiome, and to discover organisms that were not previously known.
Limitation of Next-Generation Sequencing
There are a few limitations of NGS, which are mentioned below:
- It can be costly as it requires sophisticated bioinformatics systems, fast data processing, and large data storage capabilities.
- PCR amplification prior to sequencing, may lead to PCR biases during library preparation (sequence GC-content, fragment length, and false diversity) and analysis (base errors/favoring certain sequences over others).
How to Remember
The one idea: read a million at once. Sanger reads one fragment at a time. NGS reads millions at the same time. Everything about NGS, the flow cell, the clusters, the parallel imaging, exists to make many reads happen together. Massively parallel is the whole story.
Why clusters exist: one copy is too quiet. A single DNA molecule gives a signal too faint to photograph. Bridge amplification makes about a thousand identical copies in one spot, so the cluster shouts loud enough for the camera to read. No clusters, no signal.
Sanger terminator vs NGS terminator: permanent vs reversible. Sanger's ddNTP stops the chain for good (dead end). NGS uses a reversible terminator: stop, read the base, unblock, continue. One base per cycle, then carry on. That reversibility is what lets NGS read a strand base by base instead of all at once.
The four steps in order. Library (fragment and tag) then Cluster (bridge amplify) then Sequence (synthesis, one base per cycle) then Analyze (align the reads). Library, Cluster, Sequence, Analyze.
Key exam facts in one table
| Point | Fact |
|---|---|
| Also called | Massively parallel sequencing; high-throughput sequencing |
| Core principle | Sequence millions to billions of DNA fragments at the same time |
| Dominant platform | Illumina (sequencing by synthesis) |
| Step 1 | Library preparation: fragment DNA, attach adapters (and barcodes) |
| Step 2 | Cluster generation by bridge amplification (~1,000 copies per cluster) |
| Why clusters are needed | A single molecule is too faint to detect; a cluster gives a strong signal |
| Step 3 | Sequencing by synthesis: one fluorescent, reversibly terminated base per cycle |
| Step 4 | Data analysis: align or assemble millions of short reads |
| Reversible terminator | Blocks the strand after one base, then is removed so the next base can add |
| Read type | Short reads (tens to a few hundred bases); often paired-end |
| Versus Sanger (scale) | Sanger reads one fragment per capillary; NGS reads millions at once |
| Versus Sanger (chemistry) | Sanger uses irreversible ddNTPs; Illumina uses reversible terminators |
| Main uses | Whole genomes, gene panels, RNA-seq, pathogen and outbreak detection, cancer profiling, metagenomics |
Where Students Get Confused
"What actually makes NGS 'next-generation' compared with Sanger?" Scale. Sanger reads one DNA fragment at a time. NGS reads millions to billions of fragments simultaneously on one flow cell. The massive parallelism is the defining feature and the reason NGS is so fast and cheap per base.
"Why do you need to make clusters? Why not read a single molecule?" Because a single DNA molecule gives a signal too weak for the camera to detect reliably. Bridge amplification copies each fragment about a thousand times in one tight spot, so the cluster emits a strong, clear signal when imaged. Each cluster then produces one read.
"How is the NGS terminator different from the Sanger terminator?" Sanger uses dideoxynucleotides that stop the strand permanently, so a given strand is read only up to where it stopped. NGS uses reversible terminators: the strand is blocked after each single base so it can be imaged, then the block is removed so the same strand continues to the next base. Sanger stops for good; NGS stops, reads, and resumes.
"What are adapters and barcodes for?" Adapters are short known sequences added to both ends of every fragment. They let the fragments bind the flow cell and give the machine a defined starting point. Barcodes (indexes) are short tags that mark which sample a fragment came from, so many samples can be sequenced together and then separated by computer afterward.
"What does 'sequencing by synthesis' mean?" It means the sequence is read while a new complementary strand is being built, one base at a time. Each added base carries a color that identifies it, so reading the colors in order across the cycles gives the sequence. This is the same general idea as ordinary DNA synthesis, but with a detectable, reversibly blocked base added at each step.
"Is NGS more accurate than Sanger?" Not per read. Sanger is more accurate for a single target, which is why it is still used to confirm individual findings. NGS makes up for a slightly higher per-read error rate with enormous volume and by reading each position many times (coverage), and it is unmatched for sequencing large amounts of DNA at once.
References
- Metzker M.L. (2010). Sequencing technologies: the next generation. Nature Reviews Genetics, 11(1), 31–46. https://doi.org/10.1038/nrg2626
- Goodwin S., McPherson J.D., McCombie W.R. (2016). Coming of age: ten years of next-generation sequencing technologies. Nature Reviews Genetics, 17(6), 333–351. https://doi.org/10.1038/nrg.2016.49
- Bentley D.R., Balasubramanian S., Swerdlow H.P., et al. (2008). Accurate whole human genome sequencing using reversible terminator chemistry. Nature, 456(7218), 53–59. https://doi.org/10.1038/nature07517
- Brown T.A. (2018). Genomes (use the site reference spine edition). Garland Science / CRC Press.
- Illumina. An Introduction to Next-Generation Sequencing Technology. https://www.illumina.com/science/technology/next-generation-sequencing.html
- Hilt E.E., Ferrieri P. (2022). Next-generation and other sequencing technologies in diagnostic microbiology and infectious diseases. Genes, 13(9), 1566. https://doi.org/10.3390/genes13091566
Frequently Asked Questions
What is next-generation sequencing (NGS)?
What is next-generation sequencing (NGS)?
Next-generation sequencing is a method that reads millions to billions of DNA fragments at the same time. Because so many fragments are read in parallel, it can sequence an entire genome quickly and at low cost. It is also called massively parallel sequencing or high-throughput sequencing.
What is the principle of NGS?
What is the principle of NGS?
The principle is massive parallelism. The genome is broken into millions of small fragments, each fragment is copied into a signal-strong cluster, and all the clusters are sequenced at the same time. A computer then assembles the millions of short reads back into the full sequence. This parallel reading is what separates NGS from older methods that read one fragment at a time.
What are the steps of next-generation sequencing?
What are the steps of next-generation sequencing?
The Illumina workflow has four main steps. First, library preparation: the DNA is fragmented and short adapters are attached to both ends. Second, cluster generation: the fragments bind a flow cell and are copied about a thousand times each by bridge amplification, forming clusters. Third, sequencing by synthesis: the machine reads all clusters at once, adding one fluorescent, reversibly blocked base per cycle and photographing the flow cell after each base. Fourth, data analysis: software aligns or assembles the millions of short reads into the full sequence.
What is bridge amplification?
What is bridge amplification?
Bridge amplification is how each DNA fragment is copied on the flow cell. The anchored single strand bends over and its free end binds a nearby surface oligonucleotide, forming a bridge shape. A polymerase copies across the bridge, and the strands are then separated. Repeating this makes about a thousand identical copies in one tiny spot, called a cluster. Clusters are needed because a single molecule is too faint to detect, while a thousand copies give a signal strong enough to image.
What is sequencing by synthesis?
What is sequencing by synthesis?
Sequencing by synthesis means the sequence is read while a new complementary strand is being built. In each cycle, one nucleotide is added to every cluster. Each nucleotide carries a color that identifies its base and a reversible block that stops the strand after one base. A camera records the color at every cluster, then the block and dye are removed so the next base can be added. Reading the colors in order across cycles spells out the sequence.
How is NGS different from Sanger sequencing?
How is NGS different from Sanger sequencing?
The main difference is scale. Sanger reads one fragment per capillary, while NGS reads millions to billions of fragments at once. There is also a chemical difference: Sanger uses dideoxynucleotides that stop a strand permanently, while Illumina uses reversible terminators that stop the strand after each base and are then removed so the strand continues. Sanger is more accurate for a single target and is still used to confirm results, while NGS is ideal for sequencing large amounts of DNA quickly.
What are adapters and barcodes in NGS?
What are adapters and barcodes in NGS?
Adapters are short known sequences attached to both ends of every DNA fragment. They let the fragments bind the flow cell and give the machine a defined starting point for reading. Barcodes, also called indexes, are short tags that identify which sample a fragment came from, so several samples can be sequenced together in one run and then separated by computer afterward.
What is a reversible terminator?
What is a reversible terminator?
A reversible terminator is a chemical block on a nucleotide that stops the DNA strand after just one base is added, so that base can be imaged. Unlike the permanent chain terminators used in Sanger sequencing, this block can be removed, which frees the strand to accept the next base in the following cycle. This is what allows NGS to read a strand one base at a time.
What is NGS used for?
What is NGS used for?
NGS is used to sequence whole genomes, to read targeted gene panels and whole exomes, to study gene expression through RNA sequencing, to detect and identify pathogens and track outbreaks, to profile cancer mutations, to study all the microbes in a sample without culture (metagenomics), and to detect antimicrobial resistance genes.
Is NGS more accurate than Sanger sequencing?
Is NGS more accurate than Sanger sequencing?
Not for a single read. Sanger is more accurate per read and is still used to confirm individual findings. NGS compensates with volume: it reads each position many times, and this depth of coverage, combined with computer analysis, gives reliable results across very large amounts of DNA.

Tankeshwar Acharya, MSc (Medical Microbiology)
Tankeshwar Acharya is an Assistant Professor in the Department of Microbiology at Patan Academy of Health Sciences (PAHS), Nepal, where he has been teaching and practicing clinical microbiology for over 14 years. He is the founder of Microbe Online, one of the leading free microbiology education resources on the web, covering bacteriology, mycology, parasitology, immunology, and clinical laboratory diagnostics written from direct experience in both the classroom and the diagnostic laboratory.
Comments
No comments yet. Be the first to share your thoughts.
Leave a comment
All comments are reviewed before they appear.