For research only

Research Literacy

How to Read Research Statistics in Context

Published: September 18, 2026

By the Curo Science Blog Team

A statistic is not just a number. It is a numerator, denominator, unit of analysis, selection process, method, and date. Remove any one of those and a precise figure can become difficult to interpret.

This matters in peptide research because database counts, registry totals, laboratory reports, and literature reviews are often presented beside one another even though they count different objects.

A database count begins with the query

A PubMed search returns matching records, not necessarily distinct experiments. NCBI's ESearch documentation states that the service provides identifiers matching the submitted term and supports field tags, date limits, and other query controls.

A broad synonym can retrieve papers that merely mention a term, while a narrow name can miss alternate spellings, salts, fragments, or older terminology. Reviews, editorials, animal studies, methods papers, corrections, and human trials can all appear in one result set unless the query distinguishes them.

A reproducible literature count therefore needs the database, exact query, filters, date run, and deduplication rule. The count is a property of that search, not a timeless measurement of how much evidence exists.

A registry record is not necessarily a completed trial

ClinicalTrials.gov records can describe observational or interventional studies at different phases and statuses. The ClinicalTrials.gov data API makes structured records available, but the fields still have to be interpreted separately.

Count registered studies, studies that administered the material of interest, completed studies, and records with posted results as different statistics. A mention in eligibility criteria or background text should not be counted as an intervention. A completed record without results does not supply an outcome.

Registry counts also change as records are added and updated. Preserve the query and retrieval date, then describe the result as a dated snapshot.

Define the denominator before reading a percentage

A percentage requires the population it describes. In laboratory datasets, the denominator might be submitted samples, unique lots, reports, analytes, companies, or tests. Those units produce different interpretations.

If one lot is tested repeatedly, a test-level pass rate gives that lot more weight. If only disputed or high-risk samples are submitted, the result may describe the submitted set without representing the wider market. If one report contains several assays, report count and assay count are not interchangeable.

Write the statistic as a fraction in words before interpreting it: passing assays among all assays run, unique lots meeting a criterion among lots sampled, or companies with at least one public report among companies reviewed.

Selection determines what can be generalized

Convenience samples, voluntary submissions, enforcement samples, and randomized purchases answer different questions. A public testing archive can document the material in that archive while remaining unable to estimate prevalence outside it.

Look for inclusion and exclusion criteria, missing records, repeated observations, and who chose the samples. If the selection mechanism is unknown, treat the dataset as descriptive rather than representative.

Method and threshold belong beside the result

A result depends on the procedure, measurement uncertainty, specification, and decision rule. FDA's guidance on analytical procedures and methods validation discusses characteristics such as specificity, accuracy, precision, detection limits, and robustness because methods define what measurements can support.

A pass rate assembled from unlike methods or thresholds may combine decisions that are not comparable. Preserve the assay, unit, limit, and specification for each group before calculating an aggregate.

Transparent reporting makes statistics auditable

The STARD 2015 statement defines 30 essential reporting items for diagnostic-accuracy studies. Its specific scope is medical-test accuracy, but its underlying discipline is broadly useful: readers need enough information about design, participants or samples, methods, analysis, and flow to judge bias and applicability.

For a laboratory or literature statistic, publish the query or sampling rule, the raw numerator and denominator, exclusions, grouping logic, method, date, and uncertainty. A chart without those fields can communicate scale while hiding how the scale was constructed.

Do not add statistics from different windows

Counts from overlapping time periods cannot simply be summed. Neither can a cumulative count and an annual count, or approvals and trial registrations. First align the object counted, geography, date range, and inclusion rule.

The same caution applies to updates. Replacing an older snapshot with a new one is different from adding the two snapshots together. Keep the earlier figure only when the change over time is itself the subject.

A context checklist

Before repeating a research statistic, record:

  • What is counted and what is excluded.
  • Numerator, denominator, and unit of analysis.
  • Database query or sample-selection rule.
  • Method, threshold, and grouping logic.
  • Date range and retrieval date.
  • Missing data, repeated observations, and uncertainty.
  • The population to which the result can reasonably apply.

A smaller statistic with a visible denominator is usually more informative than a larger figure whose construction cannot be inspected.