Filters: stack with AND. Click Reset to clear.
- Tissue / Subtype: DepMap OncotreeLineage and primary disease subtype.
- Sex: annotation and expression are two independent axes (see below).
- Hotspot filter: only show cell lines mutated in this gene at a hotspot position. The dropdown lists genes with live n= mutated counts (matching the Fusion / CN filter style).
- Fusion filter: the dropdown lists clinically relevant fusion pairs first (★ BCR-ABL1, ★ EWSR1-FLI1, …) followed by the gene-based long tail. Selecting a pair filters to the curated set; selecting a gene returns any line with any fusion involving that gene.
- CN amp/del filter: the dropdown lists curated focal amplifications (▲) and tumor-suppressor deletions (▼) with live n=; selecting one filters to lines carrying that focal copy-number event.
- Disease: searchable, and scoped by whatever else is filtered, so choosing Skin leaves only the skin diseases. Abbreviations work: type CLL, AML, GBM, PDAC.
- Alteration grids (▦ beside each alteration box): a grid of the cell lines currently shown, one for hotspot mutations, one for curated driver fusions, one for copy number. Click a gene in the grid to require it altered or wild-type.
- Quick filters: curated cell-line collections to include or exclude, among them the two provenance sets below.
- Interferon score (Sort: Interferon score, and the interferon-high / interferon-low quick filters): the average of 34 interferon-stimulated genes, each expressed as how far the line sits from the panel average for that gene. 0 is typical, above +0.5 is called interferon-high here and below −0.5 interferon-low. It measures the cell's own interferon signalling, which shapes how visible a line is to the immune system, and is the readout for viral mimicry. A cell line has no immune infiltrate, so it says nothing about the interferon environment of a tumor. Lines with no expression data are unscored and sort to the end. The score also appears on each cell line's own card.
- Retroelement signal (Sort: Retroelement signal, and the retroelement-high quick filter): RNA output of 750 full-length LINE-1, HERV-K and SVA elements outside genes, in counts per million, measured from the public CCLE sequencing files. The panel median is about 40 CPM; the top tenth, about 80 CPM and up, counts as retroelement-high. Lines with a higher signal depend measurably more on ADAR1, beyond what the interferon score predicts. 669 of 1,208 lines have public data; the rest are unscored and sort to the end. The score also appears on each cell line's card and wiki.
- Heatmap: draws a gene set across the cell lines the filters leave showing. Open it from the Heatmap button in this header, or from Options / Other / Gene set heatmap. Pick a preset or paste your own genes, and choose mRNA or gene effect. The column order is carried by the annotation rows under the grid: mark one row as the blocks and the colour strip, legend and Min n follow it, turn on other rows' sort arrows to order the lines within each block, and the score (the mean of the shown genes) settles ties. Rows can show hotspot mutations, fusions, copy number, lineage, subtype, disease, one gene's expression or gene effect, cell-line clustering, or the gates. Saves as an image, to the clipboard, as a .csv or as a reopenable view file, and its Methods button writes out exactly how the picture was made.
⚠ Warning marks: a triangle beside a cell line's name means Cellosaurus records a problem with it. 56 of the 1,208 lines are flagged. Hover the mark for the full explanation, or open the line to see it with a link to the source.
🦠 Virus marks: a virus beside a cell line's name means Cellosaurus records which virus the line was transformed by. 53 of the 1,208 lines carry one, most often EBV, HPV16/18 or HBV. It matters because viral oncoproteins override host pathways: an HPV-transformed line behaves as p53- and RB-deficient whatever its TP53 and RB1 sequence says, and an EBV-immortalised lymphoblastoid line is not a tumour line at all. A line without the mark simply has no such record, which is not the same as being free of the virus. The matching quick filters are named “Confirmed …-transformed” for that reason.
- Identity disputed (36 lines): the line is contaminated, misidentified, or a derivative of another line, so a result attributed to it may belong to something else. MKN28 is a MKN74 derivative; KP-1N is a PANC-1 derivative.
- Cancer type disputed (4 lines): the line is itself, but its recorded disease or tissue is genuinely unresolved, so any tissue-grouped analysis may place it incorrectly.
- Earlier classification corrected (16 lines): the type shown here is already the corrected one, so this is background rather than a warning and carries no marker in the list. SK-N-MC was called a neuroblastoma for years and is an Ewing-family tumor; KE-97 was filed as gastric and is a B-lymphoblastoid line.
- Both sets are available as quick filters, to exclude or to inspect. Nothing is hidden: a flagged line can still be a good experimental model, it just cannot be evidence about the tumor type on its label. Calls that Cellosaurus states as possible rather than settled say so.
Selecting cell lines: tick lines (or Select Visible) to build a set for Inspect / Export. Show only my N collapses the list to just your ticked lines and carries the count; it is greyed out until something is ticked, and turns off on Deselect All / Reset. The list fills column by column (down the first column in sort order, then the next).
Inspect selected vs rest: compares your selection with the cell lines it is measured against, CRISPR gene effect on the left, mRNA expression on the right. Both columns rank by the size of the difference in either direction, so genes higher and lower in your selection appear together. Each gene carries a Welch's t-test adjusted across all genes tested (Benjamini-Hochberg), shown as q and filterable. The comparison group is yours to choose: all other cell lines, only those sharing a lineage with your selection (taking tissue of origin out of the comparison), a group picked from the tissue table, or a pasted list of names. The two measures cover different cell lines, so each states its own group sizes. Either ranked gene list can be opened in the heatmap (top 10 to 100, or any number), carried over with the inspected cell lines.
Send to another view: carries the listed cell lines, or just the ticked ones, into the correlation scatter, the gene effect charts or the heatmap, either marking them there or narrowing the view to them, and can set them as the cohort for the mutation and gene set analyses. Each target row explains itself in the popout.
Detail card: clicking a line opens a card that now leads with a plain-language executive summary (cancer type, patient origin, and standout molecular features), the same overview shown when you hover a cell-line dot on any plot.
Hotspot variants: when DepMap's inferred-subtype pipeline calls a specific codon (KRAS p.G12D, BRAF p.V600E, EGFR p.L858R, JAK2 p.V617F, EGFR exon-19 del), the variant is appended in green next to the gene name in the Hotspot Mutations list. Generic group calls like KRAS p.G12 (any G12 not specifically D or C) are also shown when no specific call is available. Source: OmicsInferredMolecularSubtypes.csv.
Functional loss: integrated tumor-suppressor inactivation calls (RB1, TP53, PTEN, NF1, CDKN2A, VHL, MTAP, APC) shown as red-bordered chips in the detail panel. A gene is "lost" if any of: WGS-based relative CN < 0.3, likely-LoF mutation with AF > 0.5, or expression < 0.1 log-TPM. This catches deletion-driven losses that the damaging-mutation matrix alone misses: e.g. CDKN2A homozygous deletion in 445 lines (~37% of the panel), most of which look mutation-clean. Source: OmicsInferredMolecularSubtypes.csv.
MSI badge: small red MSI chip next to the cell-line ID when MSIsensor2 score ≥ 20 (DepMap's threshold for microsatellite instability).
Genome signatures (small row near the top of the detail panel): Ploidy (PureCN avg chromosome dosage), WGD (whole-genome doubling, shown in red when positive, ~58% of the panel), Aneuploidy (Ben-David 2021 score, 0–39), CIN (chromosomal instability). These shape interpretation when comparing gene effect across cell lines: WGD-positive lines are systematically different. Source: OmicsGlobalSignatures.csv.
Each new section in the detail panel has an inline ? with a hover tooltip recapping the key explanation, so you don't need to come back to this modal to remember what a tier or column means.
Clinically relevant fusions: shown as green/amber/red chips in the cell line detail panel.
- Curated list of ~50 well-validated driver fusions (BCR-ABL1, EWSR1-FLI1, EML4-ALK, PML-RARA, TMPRSS2-ERG, …) sourced from COSMIC fusion gene census.
- Each call is validated per cell line on three orthogonal signals: lineage match, partner expression z-score vs lineage-matched negatives, and partner dependency (CRISPR gene effect) z-score. Only the signals the curator marked as informative for that fusion are scored: e.g. BCR-ABL1 doesn't elevate ABL1 mRNA, so expression isn't checked and the call rests on the strong ABL1 dependency.
- high: both informative signals support the call (or one strong + the other directionally supportive).
- med: partial orthogonal support, or curated + lineage match without orthogonal validation.
- low: neither signal supports and lineage doesn't match the canonical disease (suspect, often artifact).
- ⚠ atypical: the cell line's tissue isn't the canonical context for this fusion (e.g. EML4-ALK in a Bowel line). Lineage is a modifier, not a gate: when orthogonal evidence is strong the call is kept and flagged, not hidden (could be a misclassified line or a genuine rare event worth investigating).
- Source: Arriba-annotated fusion calls from
OmicsFusionFilteredSupplementary.csv, filtered to confidence=high with read-through transcripts dropped.
Fusion calls, and what “not fused” means: a fusion is called from RNA-seq, so a cell line that was never sequenced is unknown, not fusion-negative. Those lines are held out of both groups rather than counted as wild-type, and the count is reported wherever a fusion comparison is made. Where the answer is published it is filled in from Cellosaurus instead: for EWSR1-FLI1 in Ewing sarcoma that returns nine cell lines, taking the fused group from 13 to 22.
Published calls also add recurrent driver fusions the caller did not report, among them EML4-ALK in NCI-H3122, TPM3-NTRK1 in KM12 and KMT2A-AFF1 in RS4;11. Private one-off fusions and ordinary immunoglobulin rearrangements are deliberately left out. Note that a line without the fusion you filtered on may carry a different one of the same gene: every EWSR1-FLI1-negative Ewing line carries EWSR1-ERG or EWSR1-FEV.
Sex: two independent axes:
- Sex (annotation): from DepMap's Model table (patient-reported). Values: Male, Female, Unknown.
- Sex (expression): computed from Y-chromosome markers (RPS4Y1, DDX3Y, EIF1AY, KDM5D, UTY, USP9Y) and XIST expression. Rules: Y-mean > 1.0 log-TPM → male; XIST > 1.0 with Y low → female; otherwise unknown. Validated on DepMap-labeled cells: ~3% false-positive on males, ~0.14% on females.
When expression is “Unknown”:
- Annotation Male → “Unknown (likely Y-chromosome loss)”: Y loss is common in cancer, esp. older male patients and advanced tumors.
- Annotation Female → “Unknown (likely XIST silencing)”: XIST silencing is documented in breast and other cancers.
Sex symbols in the list
- ♂ blue, upright: Male by annotation.
- ♀ pink, upright: Female by annotation.
- ♂ blue, italic: Male by expression only (annotation unknown).
- ♀ pink, italic: Female by expression only (annotation unknown).
- ? gray: Unknown in both axes.
Sort options, six in all. The arrow beside the picker reverses the order; for a gene or compound sort it waits until you have entered one.
- Name: alphabetical.
- GE for gene(s): type one or more genes; lines are ordered by CRISPR gene effect (Chronos), most essential first. Several genes are averaged, which ranks lines by combined essentiality.
- Expression of gene(s): type one or more genes; lines are ordered by log₂(TPM+1), highest first. Several genes are averaged, so a panel like
ASCL1, NEUROD1, CHGA, SYP finds the strongest combined signature.
- Drug response (PRISM AUC): type a compound; the picker lists PRISM Repurposing compounds ranked by how many of the visible lines they kill. Ordered by area under the curve, most sensitive first, so a lower AUC means more killing.
- Interferon score and Retroelement signal: the two scores described above, highest first; the retroelement sort adds a picker for which measurement to use (total signal, LINE-1, HERV-K, SVA, or active elements). Unscored lines sort to the end.
Gene-sort input: comma, semicolon, newline or whitespace can separate multiple symbols.
Data sources: DepMap 26Q1 for every measurement. Cell-line identity warnings and the published fusion calls come from Cellosaurus; the curated fusion and subtype panels draw on COSMIC, OncoKB, WHO and NCCN. All filters and sorts use the data currently loaded in the browser; changing tissue / subtype filters narrows the visible set that subsequent sorts operate on.