From sample intake to exportable dataset: how the pipeline works
A walk through what actually happens between logging a sample and handing an assessor a dataset — workflows, bulk import, file attachments, and generated extracts.
A sample doesn’t arrive in a laboratory as a database row. It arrives as a tube, a swab, a plate — sometimes one at a time, sometimes as a spreadsheet of two hundred of them from a field team. What happens between that arrival and an assessor being handed a clean, exportable dataset is the part most ELN write-ups skip over. This is a walk through how it actually works in Archon: workflows, bulk import, file attachments, and the datasets generated from all of it.
Start with a workflow, not a spreadsheet
Every sample moves through a workflow — a configurable sequence of stages, such as intake, processing, QC and storage, defined once per lab rather than hard-coded into the product. Each stage can require its own data capture before a sample is allowed to move to the next one, so a workflow isn’t just a label on a sample — it’s a gate. A sample can’t silently skip QC because someone forgot a field; the stage transition simply won’t complete until the data it requires exists.
This matters more than it sounds like it should, because it’s what turns “we track our samples” into something an assessor can verify rather than take on faith: the record of which stage a sample was in, and when, is the chain-of-custody history itself, not a separate document someone maintains alongside it.
Importing samples without retyping them
Field teams and instruments produce spreadsheets, not forms filled in one row at a time, so bulk import from CSV or Excel is a first-class path rather than an afterthought. It runs in two steps on purpose:
- Analyze, without writing anything. Upload the file and Archon parses it, infers a type for each column, and proposes both a column mapping and a workflow with stages that fit the data — a starting point, not a guess you’re stuck with.
- Review and correct, then import. The proposed mapping is shown before anything is written, so a column the inference got wrong gets fixed by a person who can see it, not silently imported as-is. Only the second step actually creates samples.
A single import tops out at a few thousand rows — enough for a real batch from the field, small enough that a mistake in the mapping is caught long before it becomes a few thousand mislabeled samples.
Attaching what a spreadsheet can’t hold
A row in a spreadsheet can’t hold a chromatogram, a sequencing read or a scan. Files attach directly to the sample stage they belong to — or to the notebook entry documenting the work — and the accepted formats go well past office documents and images: instrument and bioinformatics formats like FASTQ, BAM, VCF, AB1 trace files, mzML and PDB structures are all recognized natively, alongside the usual PDFs, spreadsheets and photos. Anything server- or browser-executable — scripts, HTML, binaries — is rejected outright, and every upload is written to storage under a path scoped to the organization, under a filename nobody could guess or steer. Like everything else in the pipeline, the attachment itself is an audited event, not a side effect nobody can account for later.
Turning records into a dataset
The point of collecting all of this cleanly is that it can be pulled back out cleanly. Dataset generation starts from a source — a project, a workflow, or a sample folder — then lets you apply filters and choose which columns matter, with a live preview so you can see the shape of the extract before committing to it. The result exports as CSV, Excel or JSON.
The detail that matters more than the export button: Archon stores the source, filters and columns that produced each dataset, not just the resulting file. That means the same extraction can be re-run later and shown to match — which is a very different claim from “here’s a spreadsheet someone assembled by hand in March, trust that it’s still accurate.”
None of this is a separate export tool bolted on afterward. Workflows shape how samples are collected, imports get them in without hand-typing, attachments hold what a spreadsheet can’t, and datasets are what comes back out — all of it feeding the same tamper-evident audit chain we wrote about in what makes an audit trail actually tamper-evident.
The full module-by-module walkthrough, including how this fits alongside protocols and compliance readiness, is in the lab workflow guide. If you’re comparing this against what you’d have to build or maintain yourself, the pricing page is the place to see what’s free on every tier versus what’s part of the paid ones.