Structured records and time series
Create tabular records with defined fields, relationships and constraints. Useful for software testing, scenario exploration or augmentation when the approach is supported by downstream evaluation.
Create datasets for testing, model training or gaps in coverage. We define what the data needs to represent, generate it to a specification and check it against the intended use.
Discuss a datasetCreate tabular records with defined fields, relationships and constraints. Useful for software testing, scenario exploration or augmentation when the approach is supported by downstream evaluation.
Build task-specific text examples, classification labels, extraction pairs or evaluation cases. Review consistency, duplication, coverage and the quality of labels.
Scope image generation, augmentation or rendered examples around the visual conditions that matter. Define how annotations and variations will be checked.
Produce invalid inputs, boundary conditions or less common scenarios to exercise a system. Each test case should have a clear purpose and expected behaviour.
Illustrative retail test records. The useful part is the specification behind them: valid ranges, consistent relationships and intentional test cases.
| record_id | category | quantity | unit_price | order_total | scenario |
|---|---|---|---|---|---|
| SYN-001 | home | 2 | 450.00 | 900.00 | standard_order |
| SYN-002 | electronics | 1 | 1299.00 | 1299.00 | single_item |
| SYN-003 | office | 100 | 25.00 | 2500.00 | bulk_order |
| SYN-004 | home | 0 | 450.00 | 0.00 | zero_quantity_boundary |
Example constraints: non-negative quantities, known categories, unique IDs and order_total = quantity × unit_price. Actual project rules are agreed in the dataset specification.
Validate types, ranges, required fields, relationships and domain rules. Label invalid test records deliberately so they are not confused with training data.
Check the combinations and cases specified in the brief. Inspect duplicates, class balance and unintended patterns introduced by the generator.
Where applicable, compare downstream model performance or software test coverage with an appropriate baseline. A dataset that looks plausible may still be unhelpful.
Discuss source-data permissions and potential disclosure risks. Checks may include duplication and similarity review. Formal privacy guarantees require a specific method and agreed threat model; synthetic data alone provides no such guarantee.
Agreed formats such as CSV, JSONL or image files with annotations.
Scripts or documented configuration, where included in the scope.
Checks performed, observed failures and unresolved limitations.
Schema, source assumptions, intended use and usage restrictions.
Some datasets can be generated from a schema and rules. Others need suitable source material or a simulator. We check feasibility before committing to the full dataset.
Start with the use case, the data format and the cases you need to cover.