Molecular dataset QA
Molecular dataset QA before modelling.
Check structure validity, canonical duplicates, scaffold diversity and basic molecular descriptors before a QSAR or machine-learning result starts looking more certain than the data deserves.
Quality-control scope
What we check
- 01Structure validity
We check whether submitted molecular structures can be interpreted consistently and flag rows that need review.
- 02Duplicate structures
Repeated molecular structures are surfaced so duplicate-driven leakage or accidental weighting is easier to spot.
- 03Structural diversity
We summarize scaffold-level diversity so concentrated datasets are easier to recognize before modelling.
- 04Dataset summary
The audit reports core molecular properties and quality-control counts in both human-readable and machine-readable form.
- 05Optional baseline assessment
If a target is supplied, the full audit can include a conservative baseline assessment when the data supports a meaningful evaluation.
Scope
What it does not claim
MoleculeCheck is a data-quality instrument, not a model-certification machine. A clean audit does not prove external generalization, biological relevance, clinical validity, safety, efficacy or regulatory suitability.
Paid deliverable
A reproducible audit, not a black box.
The full package contains a visual HTML report, cleaned CSV, structured JSON, a plain-text methodology report and any scientifically valid target-aware baseline.