Datasets API
Upload and manage datasets for training. These endpoints require an API key; the account is derived
from that key. X-Account-ID is optional and, if supplied, must match the key's account.
Base URL: https://api.colabhive.com/api/builder/v1
Supported direct-upload formats: csv, jsonl, parquet
See Preparing Datasets for format details and best practices.
Upload Dataset
The supported upload path is a single file request through the SDK:
from colabhive import ColabHive
client = ColabHive(
api_key="YOUR_API_KEY",
account_id="YOUR_ACCOUNT_ID",
)
dataset = client.datasets.upload(
name="my_dataset",
file="./data.csv",
)
print(dataset.dataset_id, dataset.status)
Direct Upload (REST)
POST /datasets/upload
Upload a file directly in a single multipart request.
cURL:
curl -X POST "https://api.colabhive.com/api/builder/v1/datasets/upload" \
-H "X-Account-ID: YOUR_ACCOUNT_ID" \
-H "X-API-Key: YOUR_API_KEY" \
-H "X-Dataset-Name: my_dataset" \
-H "X-Domain: text" \
-F "file=@./data.csv"
Response:
{
"success": true,
"dataset_id": "uuid",
"dataset_name": "my_dataset",
"storage_url": "s3://colabhive-datasets/datasets/.../data.csv",
"size_bytes": 245760,
"size_mb": 0.23,
"checksum": "sha256-hex",
"format": "csv"
}
The gateway accepts files up to 10 GiB. Upload capacity is bounded by the gateway and scheduler memory available during the request, so validate the intended dataset size in the target installation before relying on the upper bound.
List Datasets
GET /datasets
Query params:
limit(optional, default 20, max 100)offset(optional, default 0): number of records to skip
Python SDK:
# Offset pagination
datasets = client.datasets.list(limit=50, offset=0)
for d in datasets:
print(d.dataset_name, d.num_samples, d.status)
cURL:
curl "https://api.colabhive.com/api/builder/v1/datasets" \
-H "X-Account-ID: YOUR_ACCOUNT_ID" \
-H "X-API-Key: YOUR_API_KEY"
Get Dataset
GET /datasets/{dataset_id}
Response:
{
"dataset_id": "uuid",
"account_id": "uuid",
"dataset_name": "my_dataset",
"description": "Transaction fraud dataset",
"status": "ready",
"storage_format": "csv",
"num_samples": 5000,
"total_size_bytes": 245760,
"visibility": "private",
"created_at": "2026-01-01T00:00:00Z"
}
Python SDK:
dataset = client.datasets.get("DATASET_ID")
print(dataset.num_samples)
Delete Dataset
DELETE /datasets/{dataset_id}
Requests deletion of the dataset through the current storage service. The public contract does not yet state a physical-erasure deadline; deployments that require one must define it in writing.
Python SDK:
client.datasets.delete("DATASET_ID")
Current Storage and Upload Boundaries
The API supports the direct single-file upload above. Its contract does not claim the following controls:
- malware or antivirus scanning;
- server-side MIME sniffing independent of the declared
Content-Type; - caller-visible checksum verification;
- physical-erasure timing after
DELETE; - storage quotas and dataset versioning;
- retry idempotency for file uploads.
If a deployment requires any control above, make it an explicit deployment requirement rather than assuming it from the upload flow.
See Also
- Preparing Datasets — format details, LLM datasets, best practices
- Training API — create training runs with datasets