Skip to main content

Datasets API

Upload and manage datasets for training. These endpoints require an API key; the account is derived from that key. X-Account-ID is optional and, if supplied, must match the key's account.

Base URL: https://api.colabhive.com/api/builder/v1

Supported direct-upload formats: csv, jsonl, parquet

See Preparing Datasets for format details and best practices.


Upload Dataset​

The supported upload path is a single file request through the SDK:

from colabhive import ColabHive

client = ColabHive(
api_key="YOUR_API_KEY",
account_id="YOUR_ACCOUNT_ID",
)

dataset = client.datasets.upload(
name="my_dataset",
file="./data.csv",
)

print(dataset.dataset_id, dataset.status)

Direct Upload (REST)​

POST /datasets/upload

Upload a file directly in a single multipart request.

cURL:

curl -X POST "https://api.colabhive.com/api/builder/v1/datasets/upload" \
-H "X-Account-ID: YOUR_ACCOUNT_ID" \
-H "X-API-Key: YOUR_API_KEY" \
-H "X-Dataset-Name: my_dataset" \
-H "X-Domain: text" \
-F "file=@./data.csv"

Response:

{
"success": true,
"dataset_id": "uuid",
"dataset_name": "my_dataset",
"storage_url": "s3://colabhive-datasets/datasets/.../data.csv",
"size_bytes": 245760,
"size_mb": 0.23,
"checksum": "sha256-hex",
"format": "csv"
}

The gateway accepts files up to 10 GiB. Upload capacity is bounded by the gateway and scheduler memory available during the request, so validate the intended dataset size in the target installation before relying on the upper bound.


List Datasets​

GET /datasets

Query params:

  • limit (optional, default 20, max 100)
  • offset (optional, default 0): number of records to skip

Python SDK:

# Offset pagination
datasets = client.datasets.list(limit=50, offset=0)
for d in datasets:
print(d.dataset_name, d.num_samples, d.status)

cURL:

curl "https://api.colabhive.com/api/builder/v1/datasets" \
-H "X-Account-ID: YOUR_ACCOUNT_ID" \
-H "X-API-Key: YOUR_API_KEY"

Get Dataset​

GET /datasets/{dataset_id}

Response:

{
"dataset_id": "uuid",
"account_id": "uuid",
"dataset_name": "my_dataset",
"description": "Transaction fraud dataset",
"status": "ready",
"storage_format": "csv",
"num_samples": 5000,
"total_size_bytes": 245760,
"visibility": "private",
"created_at": "2026-01-01T00:00:00Z"
}

Python SDK:

dataset = client.datasets.get("DATASET_ID")
print(dataset.num_samples)

Delete Dataset​

DELETE /datasets/{dataset_id}

Requests deletion of the dataset through the current storage service. The public contract does not yet state a physical-erasure deadline; deployments that require one must define it in writing.

Python SDK:

client.datasets.delete("DATASET_ID")

Current Storage and Upload Boundaries​

The API supports the direct single-file upload above. Its contract does not claim the following controls:

  • malware or antivirus scanning;
  • server-side MIME sniffing independent of the declared Content-Type;
  • caller-visible checksum verification;
  • physical-erasure timing after DELETE;
  • storage quotas and dataset versioning;
  • retry idempotency for file uploads.

If a deployment requires any control above, make it an explicit deployment requirement rather than assuming it from the upload flow.


See Also​