Read the sequence. Keep the inference layers separate.
Move from sequence classes to reads, coverage, assembly, annotation, and functional comparison without treating genome size, similarity, or bulk abundance as proof.
A read is not an assembly, an assembly is not an annotation, and an annotation is not automatic functional proof. Track which inference was actually made.
01
LESSON 1 · 21 MIN
Study + retrieve
Genome organization and sequence classes
Distinguish coding, regulatory, intronic, repetitive, structural, and organelle sequence while treating function as an evidence question rather than a coding-status label.
ESSENTIAL QUESTIONWhat sequence class is described, where is it located, and what evidence supports a functional claim?
STUDY DIAGRAM · TEXT DESCRIPTION AVAILABLE
01
Do not count genes from genome size
A genome includes protein-coding genes plus introns, regulatory regions, repeated sequences, structural regions, and other noncoding DNA. Species with larger genomes do not necessarily have more protein-coding genes because sequence-class proportions differ.
Genome size includes coding and noncoding DNA
Gene count is not genome length
Repeat content can expand genome size
02
Keep sequence roles distinct
Exons retained in a mature coding RNA can contribute protein-coding sequence, whereas introns are transcribed into pre-mRNA and then removed in typical splicing. Enhancers and other regulatory DNA can affect expression without encoding protein. Some noncoding regions produce functional RNAs or contribute chromosome structure.
Introns can be transcribed then removed
Regulatory DNA can act without translation
Noncoding does not mean automatically useless
03
Require evidence for function
Conservation, expression, biochemical binding, or a perturbation can support different functional hypotheses. Coding status alone does not prove importance, and lack of a known function does not prove that a sequence is useless. Nuclear and organelle genomes are physically distinct information systems with different inheritance contexts.
Known function requires evidence
Unknown is not useless
Organelle genome ≠ nuclear chromosome
Worked example
Species A has twice the genome size of species B but similar protein-coding gene counts. What could explain the difference?
1
Genome size counts all DNA, not only coding exons.
2
Species A may contain more intronic, repeated, or other noncoding sequence.
3
The observation does not require twice as many proteins or greater organismal complexity.
ConclusionDifferent noncoding and repeated-sequence content can change genome size without a proportional change in protein-coding gene count.
Close the notes first
Retrieve the evidence boundary.
01Why is genome size not a direct gene count?
Genomes include extensive noncoding, intronic, regulatory, and repeated sequence.
Protein-coding genes are only one part of total DNA.
02Are introns absent from the primary transcript?
No; typical introns are transcribed into pre-mRNA and then removed during splicing.
Their removal is a post-transcriptional processing step.
03Does noncoding mean nonfunctional?
No; some noncoding sequences regulate, encode RNAs, or contribute structure, while other functions may be unknown.
Function is an evidence-based claim.
02
LESSON 2 · 23 MIN
Study + retrieve
Reads, assembly, alignment, and annotation
Trace raw reads through quality control, reference alignment or de novo assembly, contigs, coverage, and annotation while locating uncertainty at each stage.
ESSENTIAL QUESTIONIs this observation a raw sequence read, an assembled interval, an alignment to a reference, or a proposed biological annotation?
STUDY DIAGRAM · TEXT DESCRIPTION AVAILABLE
01
Treat reads as fragments
Sequencing produces reads from sampled DNA fragments. Quality control identifies unreliable bases or reads. Overlap among reads can support assembly into longer contigs, while reference alignment places reads against an existing sequence. Alignment and de novo assembly answer related but distinct questions.
Read = sampled fragment
Overlap can support a contig
Reference alignment uses an existing coordinate system
02
Read coverage locally
Coverage describes how many reads support positions or regions, often summarized as an average. Coverage can be uneven, so a high average does not guarantee every base is observed. Repeats longer than or too similar for the available reads can create multiple plausible placements and assembly gaps.
Average depth can hide gaps
Repeats create ambiguity
Longer context can resolve some placements
03
Keep annotation provisional
Annotation proposes genes, regulatory features, repeats, and other elements using sequence patterns, RNA or protein evidence, homology, and experiments. It is a layer of biological inference added after sequence processing, not a perfect label returned automatically by the sequencing instrument.
Sequence first, labels later
Multiple evidence types improve annotation
Annotation can be revised
Worked example
A 100-base repeat occurs in five locations, but all reads are 50 bases long and contain only repeat sequence. Why is placement ambiguous?
1
Each repeat-only read matches several genomic locations equally well.
2
The reads lack unique flanking sequence that would anchor a location.
3
A longer read or paired context reaching unique sequence could resolve some placements.
ConclusionThe sequence can be read correctly yet still be impossible to place uniquely; base calling and mapping are different uncertainties.
Close the notes first
Retrieve the evidence boundary.
01What is a contig?
A longer continuous sequence inferred by assembling overlapping reads.
Reads are fragments; overlap supports their relative order.
02Does 30× average coverage mean every base has exactly 30 reads?
No; coverage can vary and some positions may have much less or none.
An average hides local variation.
03Is annotation the same as sequencing?
No; annotation interprets the processed sequence using computational and experimental evidence.
A sequence instrument does not automatically establish feature function.
03
LESSON 3 · 22 MIN
Study + retrieve
Comparative and functional genomic evidence
Use sequence, transcript, protein, or community data to generate bounded hypotheses while separating similarity, abundance, cell composition, and causal perturbation.
ESSENTIAL QUESTIONWhat molecular layer was measured, what comparison was made, and does the result show association or a context-bounded causal effect?
STUDY DIAGRAM · TEXT DESCRIPTION AVAILABLE
01
Use similarity as a hypothesis
Sequence similarity can support shared ancestry and a possible related function, especially when combined with conserved structure or experimental evidence. It does not guarantee identical function in every organism, tissue, or environment because regulation and cellular context can differ.
Similarity supports a candidate
Context can change function
Experimental evidence strengthens assignment
02
Name the measured layer
Transcriptomics measures RNA, proteomics measures proteins, and metagenomics samples genetic material from communities. High RNA abundance does not by itself establish high active protein, and a detected microbial sequence does not alone prove which organism is metabolically active.
RNA ≠ active protein
Community DNA ≠ activity by itself
Layer determines the claim
03
Watch mixtures and causality
Bulk samples average across cell types. A bulk abundance change can reflect regulation within cells, changed cell proportions, or both. Controlled perturbation can test necessity or sufficiency in a defined context, but one edited cell line does not prove a universal organism-level mechanism.
Bulk change can be composition
Association locates candidates
Perturbation claim stays context-bounded
Worked example
Bulk tissue RNA for gene G doubles after treatment, but the tissue also contains twice as many G-high cells. What can be concluded?
1
The bulk assay averages RNA across the cell mixture.
2
A larger proportion of G-high cells can increase the bulk signal without per-cell induction.
3
Cell-type-resolved or controlled evidence is needed to separate composition from regulation.
ConclusionThe tissue-level RNA increase is real, but it does not by itself prove that treatment doubled G transcription within each cell.
Close the notes first
Retrieve the evidence boundary.
01Does the most similar sequence prove identical function?
No; it supports a functional hypothesis that still depends on context and evidence.
Related sequences can diverge in regulation or activity.
02What does transcriptomics directly measure?
RNA abundance or identity under the assay conditions.
Protein abundance and activity are downstream layers.
03Why can bulk data mislead?
A change can reflect altered cell-type composition rather than a within-cell regulatory change.
Bulk measurements average mixed populations.
Randomized retrieval set
Now locate the sequence class, processing stage, or evidence boundary.
Genome organization, repeats, coverage, assembly, annotation, transcriptomics, proteomics, metagenomics, bulk mixtures, and perturbation are interleaved.
12 PRACTICE QUESTIONS
Retrieve before you review.
Question order and all five answer options are shuffled when you begin. The correct answer stays attached to the same underlying choice.
Scope and score notice
Genomics foundations, not personal genome interpretation.
The ADA lists genomics within Genetics but does not publish a subtopic item quota. DAT TRAIN does not invent one.
Assembly algorithms, platform-specific workflows, personal genomic results, clinical variant classification, and population-genetic modeling remain outside this route unless a prompt supplies the model.
Use your results to choose what to review next—not as an official DAT score prediction.