People often lump them together as "DNA databases", but CODIS and consumer services like AncestryDNA or 23andMe are almost opposites — different molecules measured, different purposes, different rules about who can look inside. Understanding the distinction explains both how modern cases get solved and why the privacy debate is so heated.
CODIS: an identity index for investigations
CODIS — the FBI's Combined DNA Index System — is the software and the network of databases used by law-enforcement laboratories. It is built around Short Tandem Repeats (STRs): around twenty standardised, non-coding regions of DNA where a short motif repeats a variable number of times. A CODIS record is not a genome; it is a short string of numbers — the count of repeats at each location — plus a sex marker.
That design is deliberate:
- It identifies, but reveals little else. The core STR loci sit in non-coding DNA, so a CODIS profile is excellent at telling people apart but says essentially nothing about health, appearance, or behaviour.
- It is tiered and access-controlled. Local, state and national indexes (LDIS/SDIS/NDIS) feed upward, and only accredited forensic laboratories and authorised investigators can search them.
- Who is in it is defined by law. Typically: convicted offenders, arrestees (where permitted), crime-scene "forensic unknowns", and missing-persons references — not the general public.
A CODIS hit means two STR profiles match at enough loci that, in practice, they came from the same person (or an identical twin).
Consumer ancestry databases: genealogy, not policing
Services such as AncestryDNA and 23andMe are built for a completely different job — genealogy, ethnicity estimates and health traits. They don't use STRs. They use SNP microarrays, reading hundreds of thousands of single-nucleotide polymorphisms (single-letter variations) across the genome.
That richer data reveals far more than identity:
- ancestry and ethnicity estimates
- relatives, from parents down to distant cousins, by measuring shared DNA
- depending on the service, health risks, carrier status and physical traits
These databases are owned by private companies, populated by paying customers who consent to genealogical matching, and are not law-enforcement tools by default.
Where the two worlds meet: forensic genetic genealogy
The headline cases of recent years come from a bridge between them. Some third-party sites — notably GEDmatch and FamilyTreeDNA — let users upload their own raw SNP data to find relatives. Investigators realised they could convert crime-scene DNA into a compatible SNP profile, upload it, and find distant relatives of an unknown offender. Genealogists then build family trees down to a suspect, whose identity is finally confirmed with a conventional STR test.
This technique — forensic genetic genealogy — is how the "Golden State Killer" was identified in 2018. Since then, the sites involved have tightened their rules so that police matching is opt-in, reflecting the obvious tension: uploading one person's DNA can expose relatives who never chose to take a test.
The differences at a glance
| CODIS | Ancestry databases | |
|---|---|---|
| Marker | STRs (repeat counts) | SNPs (single letters) |
| Purpose | Human identification for casework | Genealogy, ethnicity, health |
| Who's in it | Offenders, arrestees, crime-scene unknowns | Paying customers who opt in |
| What it reveals | Identity + sex, little else | Ancestry, relatives, health, traits |
| Access | Accredited labs / law enforcement | The customer (and, via opt-in sites, genealogists) |
Why it matters
CODIS is narrow by design: powerful for matching a known set of people, blind to almost everything else. Ancestry databases are broad and revealing, and because relatives share DNA, a single upload implicates an entire family tree. That is exactly why forensic genetic genealogy is such an effective investigative tool — and why it sits at the centre of the modern debate about genetic privacy.