End-to-end ML pipeline

Distill AI into
lightweight inference

Use AI to build and label your training data. MLPotion learns from it and deploys a fast, deterministic scikit-learn pipeline — no GPU, no API costs, no latency at inference time.

Experiment — classification
n=9,294
EstimatorAccF1AUC
GradientBoosting 0.942 0.938 0.971
RandomForest0.9310.9270.964
LogisticRegression0.8840.8790.943
LinearSVC0.8710.8630.935
4 of 33 · 5-fold CV Deployed → /api/v1/predict/sk_…
POST /api/v1/predict/sk_7f3a… 200 12ms
{"prediction": "cost", "estimator": "GradientBoosting"}|

Pipeline

Six stages, fully integrated. Work done in each phase directly feeds the next.

01
Design

Describe what you need to the AI assistant. It proposes a dataset schema — columns, types, validation constraints — and creates it for you.

AI-assisted
02
Populate

Upload CSV, paste from a spreadsheet, ingest via API, or generate synthetic rows through the assistant. Bulk writes up to 50 rows per call.

CSV Paste API AI
03
Annotate

Consensus labelling with three independent AI annotators. Majority-vote agreement determines the label. Disagreements flagged for human review.

3-voter consensus
04
Validate

Per-column constraints enforced on every write — required fields, allowed values, numeric bounds. Training blocked until all errors resolve.

Real-time enforcement
05
Train

33 scikit-learn estimators evaluated via stratified k-fold cross-validation. Automatic text vectorisation via TF-IDF. Best pipeline selected by primary metric.

Acc F1 AUC R² RMSE
06
Deploy

REST endpoint live immediately. JSON in, prediction out. Interactive playground for testing. Pipeline artifact downloadable as .joblib.

Production-ready

Embedded assistant

Conversational interface
at every stage

An AI agent is available in every view. It can inspect your data, generate rows, set validation rules, trigger training, run predictions, and diagnose model performance.

Schema design & creation
Synthetic data generation
Consensus annotation
Validation & error resolution
Model training & diagnostics
Prediction & inference
Assistant
generate 20 rows for this dataset with realistic prompts
✓ Reading dataset schema…
✓ Added 20 rows (all passed validation)
Done. Added 20 rows covering cost, quality, and speed strategies. Dataset now has 9,314 rows, 0 validation errors.
label the unlabelled rows using consensus
✓ Labelling with AI consensus (3 voters)…
✓ Labelled 18 of 20 rows · 2 flagged for review
Labelled 18 rows confidently. 2 rows had no consensus — see the strategy_review column for the individual votes.

Integration

Standard HTTP interface

No client library required. The prediction endpoint accepts POST requests with a JSON body. A data ingestion API is also available for programmatic row insertion.

Python
import requests # predict r = requests.post( "https://<host>/api/v1/predict/<api_key>/", json={"prompt": "reduce infrastructure overhead"}, ) # ingest r = requests.post( "https://<host>/api/v1/data/<dataset_key>/", json={"prompt": "cut vendor costs", "strategy": "cost"}, )
33
Estimators
5-fold
Cross-validation
3-voter
Consensus labelling
TF-IDF
Text vectorisation
.joblib
Pipeline export