Genotyping Cyclospora: assessing current practices

Anton Nekrutenko1

1 Dept. of Biochemistry and Molecular Biology, The Pennsylvania State University, University Park, PA, USA

The current state of affairs in Cyclospora cayetanensis typing

Cyclospora cayetanensis (taxid 88456) is a single-celled parasite of the phylum Apicomplexa (taxid 5794), related to Toxoplasma (taxid 5810) and Eimeria (taxid 5800). It infects the small intestine and causes cyclosporiasis: watery diarrhoea lasting days to weeks, prone to relapse without trimethoprim-sulfamethoxazole [9].

Humans are the only host in which the parasite is known to complete its life cycle [8]. However, oocysts have been recovered from other animals as well. A systematic review of the animal literature to December 2020 found 13 relevant studies, reporting C. cayetanensis in Mediterranean mussels and carpet shell clams, in domestic and street dogs, wild chickens, wild rhesus macaques, chimpanzees and cynomolgus monkeys [16]. Its conclusion was that no naturally exposed animal has ever had its small intestine examined for endogenous stages, so no animal reservoir is confirmed, and animals shedding oocysts are more plausibly passive carriers than hosts [16]. Some of the animal reports do not survive re-inspection at all. Reanalysis of the Chinese animal literature reassigned the bovine "Cyclospora" to Eimeria subspherica (taxid 310759) on morphometric and SSU rRNA evidence, and found that several widely used reference sequences are probably E. tenella (taxid 5802) [17].

Infection comes from food or water contaminated with faeces, most often fresh produce, and the parasite has one property that determines everything downstream: the oocyst shed in faeces is not infectious. It must spend one to two weeks in the environment, sporulating, before it can infect the next person [8, 9]. There is no person-to-person transmission to trace [8].

In 2018, an unusually large year, 2,299 laboratory-confirmed cases had been reported by 1 October, more than double the previous year at the same point [1].

The 2026 season is several times that size. CDC recorded 10,468 laboratory-confirmed domestically acquired cases between 1 May and 3 August 2026, with 517 hospitalisations, 2 deaths, 47 states reporting, and more than 12,255 further cases not laboratory confirmed [18]. The health advisory that opened the season gives the baseline: 1,645 confirmed cases by 14 July, against 249 nationally over the same window in 2025 [19]. One multistate outbreak inside that total, 6,358 cases across 15 states, was traced to iceberg lettuce from Taylor Farms de Mexico in Guanajuato and recalled on 17 July [20, 21]. UKHSA reported 67 cases in returning travellers, and of the 51 with travel information, 48 had been to Mexico [22].

There is no continuous in vitro culture system, and no animal model in routine use. Experimental infection has been reported in oysters, freshwater clams, Swiss albino mice and guinea pigs [16], but nothing that the field treats as a working model. In practice almost everything known about this parasite comes from patient stool, which is also where parasite DNA is most abundant. The exceptions matter and appear later: the assay has been run on irrigation water, and a separate one was developed for fresh produce.

The genome assemblies

Forty-nine C. cayetanensis assemblies are public. None is chromosome-level, two are annotated, and the median one is in 1,391 pieces.

Distribution of contig counts and contig N50 across the 49 public Cyclospora cayetanensis assemblies
Figure 1. Assembly quality across the 49 public C. cayetanensis assemblies. None is chromosome-level and the median assembly is in 1,391 pieces.

Named points on that distribution, with three other apicomplexan references for scale:

assemblycontigscontig N50coverageannotationplatformin BRC
GCA_020976615.1 most contiguous310524 kb40×noneMiSeq + MinIONno
GCF_002999335.1 CcayRef3, the reference738193 kb20×fullMiSeqyes
GCF_000769155.1 ASM76915v23,57344 kbnot statedfull454 + GAIIxyes
GCA_002893405.1 most contigs7,910216 kb20×noneMiSeqyes
all 49, median1,391103 kb20×2 of 49MiSeq31 of 49
P. falciparum 3D7 GCF_000002765.6141.69 Mbfull
C. parvum IOWA GCF_000165345.1181.01 Mbfull
T. gondii ME49 GCF_000006565.22,5071.22 Mbfull

Three things follow. Contiguity is an order of magnitude short of the other apicomplexans, and the single most contiguous assembly—the only one of the 49 that used long reads—is one of the 18 BRC Analytics does not carry [15]. The gap originates upstream: BRC builds its catalog against UCSC's assembly hub list, which holds exactly the same 31 for this taxon [28]. Sourcing exclusively from that list is still a design choice, and closing the gap needs a supplementation route rather than only an upstream fix. Reported coverage is 20× for 35 of the 49 and no more than 40× for 44 of them. And the reference itself is a pool: its BioSample states that reads from two strains were combined to create it [24]. Where the two strains differ, that reference either collapses the difference or carries one strain's allele, so reads from the other are penalised at exactly the positions genotyping depends on.

Everything downstream is built on that. Which is why Cyclospora is genotyped from a small panel of amplicons rather than from genomes.

The CDC genotyping panel

Cyclospora is not typed one way. FDA developed a targeted amplicon sequencing assay over 52 loci, 49 of them nuclear, covering 396 known SNP sites, with an enrichment step that makes it sensitive enough for food rather than stool alone [25]. CDC and FDA have since jointly generated data on a third scheme, an expanded 63-marker panel [27]. It has 348 public runs and, as of 2026-08-12, no publication we can find, no preprint, and no released marker list: its BioSample records name the targets only as 63 markers found through Cyclospora genome, and the Methods section records how far that search went. The mitochondrial genome can also be typed on its own, which is what CFSAN's 25 runs did. More loci does not automatically mean finer resolution. Run head to head on 66 clinical specimens, the 52-locus assay resolved 24 genetic clusters against the eight-marker scheme's 27, with 15 identical between them [26]. Cluster counts alone settle nothing about which is more accurate— that needs epidemiological labels neither panel has been scored against here—but they do not support the assumption that the larger panel simply resolves more.

In the series of analyses detailed in this blog we will concentrate on the CDC's eight-marker panel because it is published, well defined, and widely used: (1) 8,522 of the 9,016 public amplicon runs used it, against 494 across the other three schemes, (2) its reference sequences, haplotype-calling coordinates and caller are all released [10], while the 63-marker coordinates are not, (3) it is the only one of the three with a published set of outbreak cluster labels to score a pipeline against [10], and (4) three groups ran it—CDC, the Public Health Agency of Canada, and FDA on irrigation water—so cross-laboratory comparison is possible.

CDC has genotyped clinical Cyclospora specimens since 2018 to support outbreak investigations. The method is targeted amplicon deep sequencing of eight markers, followed by clustering [3]. Six of the markers are nuclear and two mitochondrial [3], and they were assembled from three studies:

markersorigin
Nu_378, Nu_360i2, Mt_MSR2 nuclear, 1 mitochondrialthe genotyping scheme these were introduced with, alongside the heuristic distance measure [11]
Nu_CDS1Nu_CDS44 nucleara SNP-mining workflow across four whole genomes, covering 13 SNPs and resolving 57 stool specimens into 19 genotypes [4]
Mt-Junction1 mitochondrialthe mitochondrial junction region, evaluated on 134 laboratory-confirmed cases and yielding 14 sequence types [5]

Here is the whole panel drawn at base resolution, with primers from the three papers that introduced them [4, 5, 11]. Every base of every amplicon is shown: primer footprints in red, PART haplotype-calling windows in alternating blue with their boundaries ruled and labelled, and everything not called in grey. The nucleotides are small, enough to see the structure and be able to follow a motif, not to read comfortably at arm's length.

The four short nuclear CDS markers, two PARTs each:

Nu_CDS1 amplicon at base resolution, showing primer footprints and PART haplotype-calling windows
Figure 2. Marker Nu_CDS1. Primer footprints in red, PART windows in alternating blue, uncalled bases in grey. Gold letters beneath a site are the alternative bases observed there.
Nu_CDS4 amplicon at base resolution, showing primer footprints and PART haplotype-calling windows
Figure 3. Marker Nu_CDS4.
Nu_CDS3 amplicon at base resolution, showing primer footprints and PART haplotype-calling windows
Figure 4. Marker Nu_CDS3.
Nu_CDS2 amplicon at base resolution, showing primer footprints and PART haplotype-calling windows
Figure 5. Marker Nu_CDS2.

The two long nuclear markers, which carry most of the panel's discriminating power—16 and 20 SNPs respectively (see haplotype-sites.tsv; [11] reported 15 for Nu_378):

Nu_378 amplicon at base resolution, showing primer footprints and PART haplotype-calling windows
Figure 6. Marker Nu_378, which carries 16 of the panel's SNPs.
Nu_360i2 amplicon at base resolution, showing primer footprints and PART haplotype-calling windows
Figure 7. Marker Nu_360i2, which carries 20 of the panel's SNPs.

And the mitochondrial rRNA marker, the deepest-sequenced of the eight:

Mt_MSR amplicon at base resolution, showing primer footprints and PART haplotype-calling windows
Figure 8. Marker Mt_MSR, the mitochondrial rRNA marker and the deepest-sequenced of the eight.

The gold letters stacked beneath each site are the alternative bases seen there. Most sites are biallelic. One site carries three. The full list—marker, PART, amplicon and PART-local coordinate, reference base and observed alleles—is in haplotype-sites.tsv.

The eighth marker is not typed by alignment. This is Cmt214.A, the longest of the twenty junction references:

The Cmt214.A junction reference, showing its 5' flank, six repeat units and 3' flank
Figure 9. Cmt214.A, the longest of the twenty junction references: a 64 bp 5' flank, six 15 nt repeat units, and a 60 bp 3' flank.

Cmt214.A is 214 bp: a 64 bp 5' flank, six 15 nt repeat units spanning positions 65 to 154, and a 60 bp 3' flank. The units are not identical. Units 1 to 3 are TAGTATTATTTATAA and units 4 to 6 are TAGTATTATTTTTAA, which differ at one position [5]. Five units give a 199 bp sequence and four give 184. Types are assigned from the unit count and the motif composition, matched against the twenty reference sequences.

All twenty references, to scale:

All twenty Mt-Junction reference sequences drawn to scale
Figure 10. The twenty Mt-Junction reference sequences, drawn to scale. Types are assigned from repeat-unit count and motif composition rather than by alignment.

The sequencing data, and how much of it used this panel

Every public Cyclospora run, by project. All are short-read Illumina except one MinION run (PRJNA772675) and one 454 run (PRJNA256967).

BioProjectrunscollectedstrategyplatformpanelwho
PRJNA5789318,3252018–2025ampliconMiSeqCDC 8-markerCDC
PRJNA11304903482018–2024ampliconMiSeqCDC/FDA 63-markerCDC
PRJNA7965351862010–2021ampliconMiSeqCDC 8-markerPHAC Canada
PRJNA1052691992018–2022ampliconMiSeqFDA 52-locusFDA
PRJNA35747825not recordedampliconMiSeq, MiniSeqmitochondrion onlyCFSAN
PRJNA95255222not recordedampliconMiSeqFDA 52-locusFDA
PRJEB109121152024–2025WGSNextSeq 500UR 7510
PRJNA750933112020–2021ampliconMiSeqCDC 8-markerFDA/CFSAN, water
PRJNA437975111997–2015WGSMiSeqCDC
PRJNA14825636not recordedWGSMiSeqFDA
PRJNA77267522020WGSMiSeq, MinIONPHAC Canada
PRJNA25696722011WGA454, GAIIxCDC
PRJNA104566512016WGSMiSeqUSDA
PRJNA27955712018WGANovaSeq 6000JCVI

Collection dates come from the BioSample records, except for the two CDC projects where they are parsed from the sample names. They are year-only for almost every run. "Not recorded" means the field is missing, not available, or absent: all 22 of PRJNA952552, all 6 of PRJNA1482563, and 24 of the 25 in PRJNA357478. PRJNA1130490 has a date for 251 of its 348 runs.

9,054 runs, 966 Gbp. 9,016 are amplicon and 38 are whole-genome. Of the amplicon runs, 8,522 used the eight-marker panel: CDC's 8,325, Canada's 186, and 11 FDA water samples. The other 494 did not: 469 used the 63-marker or 52-locus panels, and 25 (PRJNA357478) were typed on the mitochondrion alone. How comparable the panels are is not yet established, because neither the 63-marker nor the 52-locus target coordinates are public.

PRJNA578931 spans specimens collected 2018 to 2025, released 2019 to 2026. Collection year and release year are different things, and a year in this post means collection, because cyclosporiasis peaks May to August and each collection year is one outbreak season.

collection year20182019202020212022202320242025
specimens1,0397979877209561,6051,316896

The table sums to 8,316; the remaining 9 of 8,325 runs have sample names that match neither parsing scheme and are excluded. The table stops at 2025. No public sequencing run of any kind carries a 2026 collection date. Searching PubMed, Europe PMC, Crossref, medRxiv and bioRxiv alongside the NCBI and ENA archives on 2026-08-11 returned thirteen items concerning the 2026 outbreak: eight agency notices [18, 19, 20, 21, 22, 29, 30, 31], three BMJ news pieces [32, 33, 34], a JAMA patient page [35], and one medRxiv preprint that analyses GenBank records from 1997 to 2022 [23]. None of them used the eight-marker panel on 2026 specimens. The preprint states the position twice, in its results and again in its limitations: "molecular surveillance data for the summer 2026 outbreak remain unavailable to the public at the time of writing" [23]. There is no produce isolate to sequence either, because no food sample has tested positive—FDA re-reviewed its one presumptive positive and withdrew it [21]. The most recent specimens on this panel are the 896 collected in 2025 and released in April 2026, which are the season before the outbreak.

Our plan

The goal is a public implementation of this workflow that runs in Galaxy, so anyone with the data can genotype Cyclospora without local infrastructure or private code. We begin with specimens generated by the same people who developed the panel and the analysis pipeline [10].

The data

203 runs inside PRJNA578931 released in November 2019, generated using MiSeq (10.3 GB in total):

BioProjectPRJNA578931
runs203 (99 Vendor A, 104 Vendor B)
collection season2018
platformIllumina MiSeq, 194 raw 2×250
panelCDC 8-marker
labels2018_gold_clusters.txt, CDC typing repository [10]
size10.3 GB

"The gold standard"

In CDC's typing workflow repository [10], in a directory called REFERENCE_CLUSTER_LIST, is a 204-line text file:

$ head -4 2018_gold_clusters.txt
Seq_ID Cluster_alias
C_IA058_18 Vendor_A
C_IA052_18 Vendor_A
C_IA040_18 Vendor_A

Two hundred and three specimens, each assigned to one of two named clusters—Vendor_A 99, Vendor_B 104. The file's provenance is traceable to the published record. Two major epidemiologically-defined cyclosporiasis clusters were identified in 2018: "one associated with salads sold by a commercial vendor (Vendor A), and the other linked to vegetable trays sold by a second vendor (Vendor B)" [1]. Vendor A accounted for 511 confirmed cases and Vendor B for 250 [2].

The labels use CDC's internal specimen ids—C_IA013_18. The specimen id lives in the BioSample Sample Alias attribute, which is a different field from SampleName and does not appear in SRA runinfo at all.

CDC's own haplotype calls for a large slice of these specimens are available. Barratt et al. present this kind of data as a barcode—one row per specimen, one column per named haplotype, and a filled box where that haplotype was detected [11]. Applied to our benchmark specimens, using CDC's own calls, it looks like this:

Barcode of CDC's deposited haplotype calls across the 153 benchmark specimens, one row per specimen and one column per named haplotype
Figure 11. CDC's deposited haplotype calls for the 153 benchmark specimens that appear in the sheet. Each box is one specimen × one named haplotype: filled means detected, white means the locus was called but that haplotype was not found, and grey across a whole locus block means no haplotype was called there at all. Locus dropout is routine, not exceptional.

Each box is one specimen × one named haplotype. Columns are grouped into the eight loci and ordered within a locus by PART, then by haplotype number. There are three states: (1) a filled box means the haplotype was detected, (2) white means the locus was called but that particular haplotype was not among the ones found, and (3) grey across a whole locus block means no haplotype was called there at all.

Where these calls come from, and how much weight they carry. They are not from a paper. The file is 2022-10-24Joel_haplotype_sheet.txt, committed to CDC's Eukaryotpying-Python repository, where the README describes it as "an example of this HDS format" [13]—deposited example input for the distance-computation code rather than a curated release of results. The data is real: 2,354 specimens, CDC's own specimen identifiers, and they join cleanly to SRA. But there is no methods statement saying which pipeline version produced these calls, no version history, and no publication that presents this sheet as its result. So it is the best available reference for what our pipeline should reproduce, and we will report concordance against it— but "CDC's deposited calls" is the honest description, not "published".

153 of the 203 appear in it, unevenly: Vendor A 98 of 99, Vendor B 55 of 104. So concordance is reported against those 153 rather than as a general accuracy figure. Mapping all 203 will show whether the absent specimens failed CDC's five-of-eight amplification criterion, which is the likeliest reason they are missing.

The median specimen carries 26.5 haplotype calls for Vendor A and 24 for Vendor B, across a median of 7 and 6 of the 8 loci respectively. Note how much grey there is: locus dropout is routine, not exceptional.

The two clusters are visibly different. Ten haplotypes differ by 50 percentage points or more between the two vendor outbreaks:

haplotypeVendor AVendor B
Nu_378_PART_D_Hap_20%95%
Nu_378_PART_D_Hap_787%0%
Mt_Cmt169.A_Junction_Hap_80%75%
Nu_CDS1_PART_B_Hap_10%67%
Nu_CDS4_PART_A_Hap_265%0%
Nu_CDS1_PART_B_Hap_264%0%
Nu_CDS1_PART_A_Hap_10%64%
Mt_Cmt199.A_Junction_Hap_1763%0%
Nu_CDS1_PART_A_Hap_257%0%
Nu_CDS4_PART_B_Hap_252%0%

Other datasets that used the same panel the same way

The eight-marker panel is not unique to CDC's own archive. Three public datasets ran it, and the BioSample records say so explicitly:

datasetrunswhowhat the record says
PRJNA5789318,325CDCGene Targets = CDS-1; CDS-2; CDS-3; CDS-4; HC378; HC360i2; Mt-Junction; MSR
PRJNA796535186Public Health Agency of Canadagene_targets = Nu_CDS1; Nu_CDS2; Nu_CDS3; Nu_CDS4; Nu_378; Nu_360i2; Mt_Cmt; Mt_MSR
PRJNA75093311FDA/CFSAN"MLST-TADS targets as described by CDC (Nascimento et al., 2020)"

8,522 runs in total. Canada uses its own naming for the same eight loci, and its BioProject description states that the markers were "originally described by the U.S. Centers for Disease Control, Accession: PRJNA578931." The 11 CFSAN runs are the only naturally contaminated non-human samples in the entire public record: surface and agricultural water from the Salinas River in California and the C-23 canal in Florida, collected between November 2020 and June 2021, with GPS coordinates and exact dates.

What we do next and what the next blog post will describe

The first thing to do with the 203 labelled specimens is to see whether an open pipeline, built only from public artifacts, recovers CDC's clusters from raw reads.

The pipeline has seven steps: (1) trim, (2) map to the seven marker references, (3) compute per-PART coverage, (4) call every haplotype at every PART, (5) assemble a specimen × haplotype presence matrix, (6) compute distances with the published heuristic [12], which CDC reported adopting in place of the earlier Bayesian ensemble [6], and (7) cluster and score against Vendor A/B.

What this design cannot tell us

Three limitations are worth stating before any results, because none of them is fixed by running the pipeline more carefully.

There is no physical ground truth. Every reference we score against is another laboratory's software output, not a specimen of known composition. A mock community—oocysts mixed at known ratios and sequenced on this panel—would separate caller error from reference error, and none exists publicly. The only laboratory-seeded material in the archive, PRJNA952552, was run on the FDA 52-locus panel rather than this one. So this gap cannot be closed with public data, and a concordance figure against CDC's calls measures agreement with CDC, not accuracy.

Differential locus dropout is an untested failure mode. The distance statistic was built to accommodate "heterogeneous (mixed) genotypes and specimens with partial genotyping data" [2], but designed-for is not the same as measured. The case that matters is two specimens that drop different loci: their similarity is then computed over whichever loci happen to survive in both. Part 2 will mask loci in silico at the dropout rates we observe and report how far the distance matrix and the cluster score move.

Joining archive metadata is more fragile than it looks. Collection year and state are parsed from sample names rather than read from structured fields, and names matching no known scheme are recorded as unknown rather than guessed. That catches parsing failures but not the worse case, where parsing succeeds against the wrong field. A concrete instance: ENA's sample_alias for PRJNA578931 returns the SampleName (18USIAxxx1056CcS0011), not the CDC specimen identifier (C_WI090_18) that the haplotype sheet is keyed on. Joining on it returns zero matches out of 8,325—a clean, confident and entirely wrong answer.

One further constraint belongs to the assay rather than to us. The eight markers are separate PCR products, so no read pair spans two of them and haplotypes cannot be phased across loci by any method. Within a PART the haplotype is already the phased unit, a whole sequence variant rather than independent SNPs. Whether two haplotypes at different loci come from one strain or from a mixed infection is therefore not recoverable from this data, for us or for anyone.

One thing we are not doing, stated plainly given the limitations above. CDC's production system is not public—its data statement directs readers to contact the authors [7]. We are not reimplementing that system and will not claim to. What we are building is an independent open implementation of the published method, benchmarked against published labels and, with the caveats above, against CDC's deposited calls. That is a narrower claim, and a checkable one.

Methods

Everything in this post that is not attributed to a source in the list below, we computed. Every figure and table derives from public URLs and APIs queried at run time rather than from a local copy, and the queries and procedures are described below.

The sequencing landscape. Run counts, library strategies, base counts and release dates come from an NCBI esearch on Cyclospora[Organism] in the sra database, fetched as runinfo via an eutils history handle, cross-checked against the ENA Portal API for the same taxon. ENA returns 8,211 runs against NCBI's 9,054, and we reconciled that gap accession by accession rather than picking the larger number. Every ENA run is present in NCBI—the ENA-only set is empty—so ENA is a strict subset and the 843-run difference is mirroring lag, accounted for by PRJNA578931 (690), PRJNA1130490 (147) and PRJNA1482563 (6) and dominated by the most recent releases (555 from April 2026, 140 from March 2026). Query scope was checked too: Cyclospora[Organism], txid88456[Organism:exp] and Cyclospora cayetanensis[Organism] return identical counts, and no run in SRA is assigned to the genus but not to this species, so the genus-level query and the species-level assembly census cover the same material. Assembly counts come from the NCBI Datasets API for taxon 88456.

Collection year and US state. Neither is in a structured BioSample field. Both are parsed out of the sample name, which uses two schemes across release batches, 19USNY13G1196CcS0011 and DC25-FL0040-Cx1-S. Names matching neither scheme are recorded as unknown rather than guessed. This is also how the year-by-year counts in the data section were produced.

Which datasets used which panel. From the Gene Targets / gene_targets attribute on the BioSample records themselves, fetched with efetch db=biosample in batches of 100. The free-text method description on the FDA water samples was read directly from those records.

What is not published about the 63-marker panel. This is a negative claim, so here is the extent of it. The run count is NCBI's: 348 read runs for PRJNA1130490. The BioProject record carries no linked publication. Europe PMC full-text search returns zero results for "63 markers" AND Cyclospora, zero for "63 genetic markers" AND Cyclospora, and zero for the accession PRJNA1130490 anywhere in full text, and none of the 23 indexed Cyclospora preprints concerns it. The two recent CDC papers that might have described it both state the eight-marker panel explicitly instead [3, 26]. For the markers themselves, CDC's typing workflow repository has 137 files, none referring to 63 markers or an expanded panel, and its reference FASTA still contains exactly seven non-junction records [10]. The BioSample attribute reads Gene Targets = 63 markers found through Cyclospora genome, with no names and no coordinates. All checked 2026-08-12. This establishes that the panel is undescribed in the indexed literature and in CDC's public code, not that no description exists anywhere.

The 203-specimen manifest. CDC's specimen identifiers appear in the BioSample Sample Alias attribute, which is a different field from SampleName and absent from SRA runinfo. Each of the 203 labels was resolved by esearch db=biosample on the identifier, then elink to the sra database, then esummary for the run accession, and finally the ENA search endpoint for the FASTQ URLs. Identifiers were normalised for case and for hyphen-versus-underscore before joining.

Read lengths. Measured directly. FASTQ files were downloaded from ENA and every read length counted (SRR10395970, SRR10415148, SRR31736733). The avgLength field in SRA runinfo is the combined length of a read pair, which is how the raw-versus-pre-trimmed split across the 203 was determined.

Primer placement and PART boundaries. Each published primer sequence was located in the marker reference FASTA by exact string match, including the reverse complement for reverse primers. We assert that every forward primer begins at base 1, that every reverse primer ends at the final base, and that the PART intervals in the BED coincide with the primer footprints. All seven markers pass, which is the basis for the claim that the PARTs are the primer-free interior of each amplicon.

Haplotype-defining sites. For each PART, the named haplotypes were compared column by column and every position where they differ recorded, together with the alleles observed there (output in haplotype-sites.tsv). PARTs with a single named haplotype, or whose haplotypes differ in length, are excluded from this count rather than aligned.

The expected-truth barcode. Built by joining CDC's deposited haplotype sheet to the manifest on the normalised specimen identifier (output in expected-truth.tsv).

Sources

  1. Barratt JLN, et al. (2021) Investigation of US Cyclospora cayetanensis outbreaks in 2019 and evaluation of an improved Cyclospora genotyping system against 2019 cyclosporiasis outbreak clusters. Epidemiology and Infection 149:e214. doi:10.1017/S0950268821002090 · PMID 34511150 · PMC8506454
  2. Nascimento FS, et al. (2020) Evaluation of an ensemble-based distance statistic for clustering MLST datasets using epidemiologically defined clusters of cyclosporiasis. Epidemiology and Infection 148:e172. doi:10.1017/S0950268820001697 · PMID 32741426 · PMC7439293
  3. Peterson A, et al. (2025) Assessing the sequencing success and analytical specificity of a targeted amplicon deep sequencing workflow for genotyping the foodborne parasite Cyclospora cayetanensis. Journal of Clinical Microbiology. doi:10.1128/jcm.01811-24 · PMID 40366167 · PMC12153321
  4. Houghton KA, et al. (2020) Development of a workflow for identification of nuclear genotyping markers for Cyclospora cayetanensis. Parasite 27:24. doi:10.1051/parasite/2020022 · PMID 32275020 · PMC7147239
  5. Nascimento FS, et al. (2019) Mitochondrial junction region as genotyping marker for Cyclospora cayetanensis. Emerging Infectious Diseases 25(7). doi:10.3201/eid2507.181447 · PMID 31211668 · PMC6590752
  6. Barratt JLN, et al. (2023) Cyclospora cayetanensis comprises at least 3 species that cause human cyclosporiasis. Parasitology 150:269–285. doi:10.1017/S003118202200172X · PMID 36560856 · PMC10090632
  7. Jacobson D, et al. (2023) Novel insights on the genetic population structure of human-infecting Cyclospora spp. Current Research in Parasitology & Vector-Borne Diseases. doi:10.1016/j.crpvbd.2023.100145 · PMID 37841306 · PMC10569985
  8. Dubey JP, Khan A, Rosenthal BM (2022) Life cycle and transmission of Cyclospora cayetanensis: knowns and unknowns. Microorganisms 10:118. doi:10.3390/microorganisms10010118 · PMID 35056567
  9. Almeria S, Cinar HN, Dubey JP (2019) Cyclospora cayetanensis and cyclosporiasis: an update. Microorganisms 7:317. doi:10.3390/microorganisms7090317 · PMID 31487898
  10. CDC Cyclospora typing workflow (alpha test release), including REFERENCE_CLUSTER_LIST/2018_gold_clusters.txt, the marker and junction reference FASTAs, and CYCLOSPORA_FEB_11_2020.bed. github.com/Joel-Barratt/CDC-Complete-Cyclospora-typing-workflow-ALPHA-TEST
  11. Barratt JLN, et al. (2019) Genotyping genetically heterogeneous Cyclospora cayetanensis infections to complement epidemiological case linkage. Parasitology 146:1275–1283. doi:10.1017/S0031182019000581 · PMID 31148531—Fig. 3 is the barcode layout reproduced here.
  12. Eukaryotyping—R implementation of Plucinski's Bayesian method and Barratt's heuristic definition of genetic distance. github.com/Joel-Barratt/Eukaryotyping
  13. Eukaryotpying-Python (repository name misspelled upstream), including haplotype_sheets/2022-10-24Joel_haplotype_sheet.txt. github.com/Joel-Barratt/Eukaryotpying-Python
  14. NCBI BioProject PRJNA578931, Cyclospora Genotyping (CDC). www.ncbi.nlm.nih.gov/bioproject/PRJNA578931
  15. BRC Analytics organism page for Cyclospora cayetanensis (taxid 88456). brc-analytics.org/data/organisms/88456
  16. Totton SC, O'Connor AM, Naganathan T, Martinez BAF, Sargeant JM (2021) A review of Cyclospora cayetanensis in animals. Zoonoses and Public Health. doi:10.1111/zph.12872 · PMID 34156154
  17. Feng K, Guo Y, Li N, Xiao L, Feng Y (2025) Cyclospora in humans, animals, fresh produce and water in China: implications for host specificity of Cyclospora species and zoonotic transmission of C. cayetanensis. One Health Advances 3:24. doi:10.1186/s44280-025-00094-y
  18. CDC, Surveillance of Cyclosporiasis, updated 2026-08-04. Archived 2026-08-09
  19. CDC Health Advisory CDCHAN-00531, Domestically Acquired Cyclosporiasis Cases in Multiple U.S. States, 2026, issued 2026-07-14. Archived 2026-08-09
  20. CDC outbreak notice, Cyclospora Outbreak Linked to Iceberg Lettuce, published 2026-07-14, updated 2026-08-05. Archived 2026-08-09
  21. FDA CORE advisory, Investigation of 15-State Outbreak of Cyclospora Illnesses: Iceberg Lettuce (July 2026), current as of 2026-08-05. Archived 2026-08-09 The live URL has been rewritten as the state count grew, so the snapshot is what the figures quoted here refer to.
  22. UK Health Security Agency, Sharp rise in cyclospora infections linked to Mexico travel, 2026-07-30. Archived 2026-08-07
  23. Janies D, et al. (2026) A map of the historical spread of Cyclospora cayetanensis with a focus on the USA. medRxiv, posted 2026-08-03. doi:10.64898/2026.08.01.26359445 · Europe PMC PPR1291175. Note the openRxiv 10.64898 prefix—the 10.1101 form of this DOI does not resolve.
  24. NCBI Assembly GCF_002999335.1 (CcayRef3), the C. cayetanensis reference genome, and its BioSample SAMN08618443, which states: "This Biosample consists of two strains of C.cayetanensis. Sequence reads from the two strains have been combined to create the Reference Genome Assembly." www.ncbi.nlm.nih.gov/biosample/SAMN08618443
  25. Leonard SR, Mammel MK, Gharizadeh B, et al. (2023) Development of a targeted amplicon sequencing method for genotyping Cyclospora cayetanensis from fresh produce and clinical samples with enhanced genomic resolution and sensitivity. Frontiers in Microbiology 14:1212863. doi:10.3389/fmicb.2023.1212863 · PMID 37396378
  26. Leonard SR, Mammel MK, Almeria S, et al. (2024) Evaluation of the increased genetic resolution and utility for source tracking of a recently developed method for genotyping Cyclospora cayetanensis. Microorganisms 12:848. doi:10.3390/microorganisms12050848 · PMID 38792677
  27. NCBI BioProject PRJNA1130490, "Cyclospora genotyping expanded panel" (CDC), whose description states that it "contains NGS data generated by CDC and the U.S. Food and Drug Administration (FDA) using an expanded Cyclospora genotyping panel comprising 63 genetic markers spread throughout the Cyclospora genome." www.ncbi.nlm.nih.gov/bioproject/PRJNA1130490
  28. UCSC BRC assembly hub list, the upstream source BRC Analytics builds its catalog against. It contains 31 C. cayetanensis assemblies, the same 31 BRC serves, and does not contain GCA_020976615.1. hgdownload.soe.ucsc.edu/hubs/BRC/assemblyList.json
  29. CDC, Investigation Update: Cyclospora Outbreak, July 2026, published 2026-07-16, updated 2026-08-05. Its "Traceback and laboratory data" section names no genotyping method. Archived 2026-08-08
  30. CDC, Cyclosporiasis Outbreaks and Investigations (index page), updated 2026-08-05. Archived 2026-08-07
  31. FDA CORE Network, Investigations of Foodborne Illness Outbreaks (tracking table). It lists seven distinct Cyclospora incidents for 2026, only one of which has a public CDC outbreak notice. Archived 2026-08-09
  32. Brown C (2026) Cyclosporiasis: as "explosive diarrhoea" sweeps US, what's behind the outbreak and what is Trump's role? BMJ, 2026-07-17. doi:10.1136/bmj-2026-100317 · PMID 42468983
  33. Wise J (2026) Cyclosporiasis: should people avoid fruit and vegetables? BMJ, 2026-07-20. doi:10.1136/bmj-2026-100324 · PMID 42476611
  34. O'Dowd A (2026) Cyclosporiasis: "explosive diarrhoea" spreads to UK in outbreak linked to Mexico. BMJ, 2026-07-31. doi:10.1136/bmj-2026-100462 · PMID 42538041
  35. Linder KA, Malani PN (2026) What is cyclosporiasis? JAMA Patient Page, 2026-07-21. doi:10.1001/jama.2026.14866 · PMID 42479478

Data: BioProject PRJNA578931 [14]. Labels, panel references and BED from CDC's public typing workflow repository [10]. Deposited haplotype calls from [13]. Genomes and annotation via the BRC Analytics organism page [15]. Every number in this post was pulled live on 2026-08-11 and can be re-derived from public sources.