Your datasets deserve version control
Git for your computer vision data. Track changes, compare versions, and ensure every experiment is reproducible.
v3
Current Version
- +2,340 images
- 24 Datasets
- 100% Reproducibility
- 0 Data loss
- 5x Faster iterations
- ∞ Version history
Used by teams at
Track every change, reproduce any result
Your datasets evolve constantly—new images, corrected labels, filtered samples. Without version control, you're flying blind.
Immutable snapshots
Every dataset version is a permanent snapshot. Reference exact data states in experiments.
Label management
Create, rename, and merge labels across your dataset. Keep your taxonomy clean and consistent.
Fork for experiments
Fork dataset versions to test hypotheses without affecting production data.
Audit-ready history
Complete changelog with who changed what, when, and why. Perfect for compliance.
Version Timeline
Interactive
- v1.0
Jan 15 - v1.1
Feb 02 - v2.0
Feb 28
production - v2.1
Mar 15
Dataset: defect-detection
- Last modified: Feb 28
- Images: 2,100
- Annotations: 8,400
- Change: +650
Everything you need to manage datasets
Version, organize, and share your datasets. Everything connects to your experiments.
Git-like Version Control
Track every change to your datasets. Compare versions, rollback mistakes, and branch for experiments. Full lineage from raw data to trained models.
- 100% Reproducibility
- 5x Faster discovery
Smart Data Organization
Tag, filter, and slice your data in seconds. Create custom views, save queries, and share collections. No more hunting through folders.
- 60% Less coordination
Team Collaboration
Share datasets across teams with fine-grained permissions. Track who changed what, when, and why. Comments and reviews built-in.
- 100% Audit compliance
Full Data Lineage
Trace any prediction back to its training data. Audit-ready lineage for compliance. Understand model behavior through data.
Programmatic dataset management
Full Python SDK with type hints, auto-completion, and comprehensive documentation. Integrate datasets directly into your ML pipelines.
Example Code
from picsellia import Client
client = Client()
datalake = client.get_datalake()
# Get or create dataset
dataset = client.get_dataset("defect-detection")
# Create a new version
version = dataset.create_version(
version="v3",
description="Added edge cases"
)
# Add data from datalake
data = datalake.list_data(
tags=["edge-case", "validated"]
)
version.add_data(data)
Dataset Browser
- GridSplit
- train: 8,400 (70%)
- validation: 1,800 (15%)
- test: 1,800 (15%)
- Total images: 12K
- Total annotations: 48K
- Balanced Class dist.
Structure your data the right way
Proper data splits are crucial for model performance. Create reproducible train/val/test splits, stratify by class, and ensure no data leakage.
- Automatic stratified splits by class distribution
- Custom split ratios with reproducible seeds
- No overlap guarantee between splits
- Re-split without losing annotations
Fits into your existing workflow
Datasets connect directly to annotations, experiments, and deployments. No manual handoffs.
Ready to version your datasets?
Free trial, no credit card. Start versioning your datasets today.
- No credit card required
- 14-day free trial
- Unlimited versions
- 50M+ Images versioned
- 100% Reproducibility
- 0 Data loss