Molecular dataset QA

Molecular dataset QA before modelling.

Check structure validity, canonical duplicates, scaffold diversity and basic molecular descriptors before a QSAR or machine-learning result starts looking more certain than the data deserves.

CSV / SMILES ≤ 20,000 rows Files scheduled for deletion after 24 h
AnalysisDeterministic checks
AccountsNone
Retention24 hours
Billing€29 once

Quality-control scope

What we check

  1. 01
    Structure validity

    We check whether submitted molecular structures can be interpreted consistently and flag rows that need review.

  2. 02
    Duplicate structures

    Repeated molecular structures are surfaced so duplicate-driven leakage or accidental weighting is easier to spot.

  3. 03
    Structural diversity

    We summarize scaffold-level diversity so concentrated datasets are easier to recognize before modelling.

  4. 04
    Dataset summary

    The audit reports core molecular properties and quality-control counts in both human-readable and machine-readable form.

  5. 05
    Optional baseline assessment

    If a target is supplied, the full audit can include a conservative baseline assessment when the data supports a meaningful evaluation.

Scope

What it does not claim

MoleculeCheck is a data-quality instrument, not a model-certification machine. A clean audit does not prove external generalization, biological relevance, clinical validity, safety, efficacy or regulatory suitability.

Paid deliverable

A reproducible audit, not a black box.

The full package contains a visual HTML report, cleaned CSV, structured JSON, a plain-text methodology report and any scientifically valid target-aware baseline.