Hyperparameter Tuning
Reference for all configurable hyperparameters per model family, with defaults, ranges, and tuning strategies.
Tip: Use
GET /training/model-configs/{model_id}/schemato get the live JSON Schema for any model's hyperparameters — always up to date.
Target column (tabular templates):
target_columnis not a declared schema parameter. The training node promotes it to the declared list parametertarget_columns("price"→["price"]), so it only takes effect on templates that declaretarget_columns; each table says which. With no target given, the target is auto-detected: the analyzed dataset schema's target, else a column namedtarget,label,outputory, else the last column.
Seeds: the tabular and BERT templates that declare
seeddefault tonull, which keeps each template's built-in behavior. Pass an integer for reproducible runs.
Tree-Based Models (XGBoost, LightGBM, CatBoost, Random Forest)
XGBoost Regression / Classification
Model IDs: xgboost-regression, xgboost-regression-gpu, xgboost-classification
| Parameter | Default | Range | Notes |
|---|---|---|---|
n_estimators | 100 | 50-5000 | xgboost-regression, xgboost-classification only. Number of trees. More = better but slower |
n_estimators | 100 | 10-1000 | xgboost-regression-gpu only. Number of trees (this ID accepts at most 1000) |
max_depth | 6 | 3-12 | Tree depth. Higher = more complex patterns, more overfitting risk |
learning_rate | 0.1 | 0.01-0.5 | Step size. Lower LR + more trees = better generalization |
subsample | 1.0 | 0.5-1.0 | Fraction of samples per tree. Values below 1.0 reduce overfitting |
colsample_bytree | 1.0 | 0.5-1.0 | Fraction of features per tree |
reg_lambda | 1 | 0-10 | xgboost-regression, xgboost-classification only. L2 regularization |
objective | null | "binary:logistic", "multi:softprob" | xgboost-classification only. Learning objective. null picks binary:logistic for a two-class target and multi:softprob otherwise. Just these two probability objectives are accepted, because the template's evaluation and the serving path read class probabilities. binary:logistic needs a two-class target; multi:softprob returns one probability per class |
eval_metric | null | "auc", "aucpr", "error", "logloss", "merror", "mlogloss" | xgboost-classification only. Metric XGBoost logs on the training data every boosting round. The template does not stop early, so the metric does not change the fitted model. null means logloss with binary:logistic and mlogloss with multi:softprob. error and logloss need binary:logistic, merror and mlogloss need multi:softprob, and auc and aucpr work with both. A mismatch with an explicit objective is refused before queuing; with objective unset, it fails before training |
min_child_weight | 1.0 | 0 or more | xgboost-classification only. Minimum sum of instance weight (hessian) needed in a child. Higher = more conservative trees |
scale_pos_weight | null | greater than 0 | xgboost-classification only. Weight of positive relative to negative examples for binary:logistic, for unbalanced classes; a typical value is the number of negatives divided by the number of positives. null keeps XGBoost's default, 1. It must be greater than 0 (0 itself is refused). Refused with multi:softprob, and a run whose target has more than two classes fails before training |
seed | null | 0-2147483647 | Random seed. null keeps the built-in behavior |
target_column | auto-detected | — | Name of the column to predict. Promoted to target_columns |
target_columns | — | list of column names | xgboost-regression, xgboost-regression-gpu only. Several target columns (multi-output) |
target_columns | null | one column name, as a one-item list | xgboost-classification only. One explicit target column; null auto-detects it. target_column is its singular alias, and conflicting values are refused before queuing |
Quick presets:
# Small dataset (<10k) — more regularization
hyperparameters={
"n_estimators": 200,
"max_depth": 4,
"learning_rate": 0.05,
"subsample": 0.8,
}
# Large dataset (>100k) — more trees
hyperparameters={
"n_estimators": 1000,
"max_depth": 6,
"learning_rate": 0.05,
"subsample": 0.8,
}
LightGBM Regression / Classification
Model IDs: lightgbm-regression-gpu, lightgbm-classification
| Parameter | Default | Range | Notes |
|---|---|---|---|
n_estimators | 100 | 100-5000 | Number of boosting rounds |
num_leaves | 31 | 20-255 | Max leaves per tree (key LightGBM param) |
max_depth | 6 | 1-20 | lightgbm-regression-gpu only. Depth cap applied on top of num_leaves |
max_depth | -1 | -1 (unlimited) to 50 | lightgbm-classification only. -1 means controlled by num_leaves |
learning_rate | 0.1 | 0.01-0.3 | |
feature_fraction | 1.0 | 0.1-1.0 | Fraction of features per tree (LightGBM's name for colsample_bytree) |
lambda_l2 | 0.0 | 0-10 | L2 regularization |
subsample | 1.0 | 0.5-1.0 | lightgbm-classification only. Fraction of samples per tree (bagging turns on below 1.0) |
seed | null | 0-2147483647 | Random seed. null keeps the built-in behavior |
target_column | auto-detected | — | Promoted to target_columns |
target_columns | — | list of column names | lightgbm-regression-gpu only. Several target columns (multi-output) |
target_columns | null | one column name, as a one-item list | lightgbm-classification only. One explicit target column; null auto-detects it. target_column is its singular alias, and conflicting values are refused before queuing |
Key difference vs XGBoost: Use num_leaves instead of max_depth. Start from the default num_leaves=31 and tune from there.
Random Forest Regression / Classification
Model IDs: random-forest-regression, random-forest-classification
| Parameter | Default | Range | Notes |
|---|---|---|---|
n_estimators | 100 | 50-500 | Number of trees |
max_depth | 10 | 1-50 | Max tree depth. Higher = deeper, more specific trees |
min_samples_split | 2 | 2-20 | Min samples to split a node |
min_samples_leaf | 1 | 1-10 | Min samples at leaf |
seed | null | 0-2147483647 | Random seed. null keeps the built-in behavior |
target_column | auto-detected | — | Promoted to target_columns |
target_columns | — | list of column names | random-forest-regression only. Several target columns (multi-output) |
target_columns | null | one column name, as a one-item list | random-forest-classification only. One explicit target column; null auto-detects it. target_column is its singular alias, and conflicting values are refused before queuing |
CatBoost Regression / Classification
Model IDs: catboost-regression, catboost-classification
| Parameter | Default | Range | Notes |
|---|---|---|---|
n_estimators | 500 | 100-5000 | Number of trees (CatBoost's iterations) |
max_depth | 6 | 4-10 | Tree depth (CatBoost's depth) |
learning_rate | 0.03 | 0.01-0.3 | |
l2_leaf_reg | 3 | 1-10 | L2 regularization |
subsample | 1.0 | 0.1-1.0 | catboost-regression only. Fraction of samples per tree |
subsample | null | greater than 0, up to 1 | catboost-classification only. Bagging sample rate; it must be greater than 0 (0 itself is refused) and at most 1. null keeps CatBoost's default bootstrap: MVS at 0.8 for a binary target on CPU with 100 rows or more, and Bayesian, which has no sample rate, for a multi-class target and on GPU. When set, a binary target on CPU keeps MVS; elsewhere the template switches to Bernoulli bootstrap. effective_parameters.json records the bootstrap type and rate used |
seed | null | 0-2147483647 | Random seed. null keeps the built-in behavior |
target_column | auto-detected | — | Promoted to target_columns |
target_columns | — | list of column names | catboost-regression only. Target column list |
target_columns | null | one column name, as a one-item list | catboost-classification only. One explicit target column; null auto-detects it. target_column is its singular alias, and conflicting values are refused before queuing |
CatBoost handles categorical features natively — no need to encode string columns before training.
Linear & Logistic Regression
Model IDs: linear-regression, logistic-regression
The two templates declare different parameters.
Linear Regression
Model ID: linear-regression
| Parameter | Default | Options | Notes |
|---|---|---|---|
regularization | "ridge" | "ridge", "lasso", "elasticnet", "none" | Type of regularization |
alpha | 1.0 | 0.001-100 | Regularization strength. Higher = stronger |
l1_ratio | 0.5 | 0-1 | ElasticNet mix (0=Ridge, 1=Lasso) |
fit_intercept | true | true, false | Fit an intercept term |
target_column | auto-detected | — | Ignored by this template: the target is auto-detected |
Logistic Regression
Model ID: logistic-regression
| Parameter | Default | Options | Notes |
|---|---|---|---|
C | 1.0 | 0.0001-1000 | Inverse regularization strength. Lower = stronger regularization |
penalty | "l2" | "l1", "l2", "elasticnet", "none" | Type of regularization |
l1_ratio | 0.5 | 0-1 | ElasticNet mix, used when penalty is "elasticnet" |
max_iter | 1000 | 100-10000 | Max solver iterations |
fit_intercept | true | true, false | Fit an intercept term |
seed | null | 0-2147483647 | Random seed. null keeps the built-in behavior |
target_column | auto-detected | — | Ignored by this template: the target is auto-detected |
SVM / SVR
Model IDs: svm-classification, svr-regression
| Parameter | Default | Options | Notes |
|---|---|---|---|
kernel | "rbf" | "rbf", "linear", "poly", "sigmoid" | Kernel type |
C | 1.0 | 0.01-1000 | Regularization. Higher = less regularized |
gamma | "scale" | "scale", "auto" | Kernel coefficient (numeric values are not accepted) |
degree | 3 | 2-10 | Polynomial degree, used by the "poly" kernel |
epsilon | 0.1 | 0-10 | svr-regression only. Width of the tube with no penalty |
seed | null | 0-2147483647 | svm-classification only. Random seed. null keeps the built-in behavior |
target_column | auto-detected | — | Ignored by these templates: the target is auto-detected |
SVM is slow on large datasets (>50k rows). Prefer XGBoost or LightGBM for larger data.
Deep Learning (MLP, TabNet)
MLP (Multi-Layer Perceptron)
Model ID: mlp-tabular-gpu
| Parameter | Default | Range | Notes |
|---|---|---|---|
hidden_dims | [128, 64, 32] | List of ints | Layer sizes |
learning_rate | 0.001 | 0.0001-0.01 | Adam LR |
batch_size | 64 | 16-512 | Training batch size |
epochs | 50 | 10-500 | Training epochs |
target_column | auto-detected | — | Promoted to target_columns |
target_columns | — | list of column names | Several target columns (multi-output) |
TabNet
Model ID: tabnet-tabular-gpu
| Parameter | Default | Range | Notes |
|---|---|---|---|
n_steps | 3 | 1-10 | Number of sequential attention steps |
n_d | 8 | 4-64 | Decision embedding dimension |
n_a | 8 | 4-64 | Attention embedding dimension |
gamma | 1.3 | 1.0-2.0 | Feature reuse across steps. Higher = more reuse allowed |
learning_rate | 0.02 | 0.001-0.1 | |
batch_size | 256 | 64-1024 | |
epochs | 100 | 50-500 | |
target_column | auto-detected | — | Promoted to target_columns |
target_columns | — | list of column names | Several target columns (multi-output) |
NLP (BERT Classification)
Model ID: bert-classification-gpu
| Parameter | Default | Range | Notes |
|---|---|---|---|
epochs | 3 | 2-20 | Fine-tuning epochs |
batch_size | 16 | 8-128 | Reduce if OOM errors |
learning_rate | 2e-5 | 1e-6 to 5e-5 | Typical BERT range |
max_length | 128 | 64-512 | Max token length per sample |
text_column | "text" | — | Column with text data |
label_column | "label" | — | Column with class labels |
base_model | "bert-base-uncased" | — | Hugging Face encoder to fine-tune |
seed | null | 0-2147483647 | Random seed. null keeps the built-in behavior |
LLM Fine-Tuning (llm-qlora-finetune)
Model ID: llm-qlora-finetune
See the LLM Fine-Tuning Guide for a full walkthrough.
| Parameter | Default | Options / Range | Notes |
|---|---|---|---|
base_model | Qwen/Qwen2.5-Coder-7B-Instruct | Any compatible HF CausalLM | HuggingFace repo ID |
adapter | "qlora" | "qlora", "lora", "none" | none = full fine-tune (needs more VRAM) |
epochs | null | 1-10 | null = 3 epochs (alias num_epochs) |
learning_rate | 0.0002 | 1e-5 to 5e-4 | Typical LoRA range |
lora_r | 16 | 4-128 | LoRA rank. Higher = more parameters |
lora_alpha | 32 | 8-256 | Usually lora_r * 2 |
lora_dropout | 0.05 | 0-0.2 | |
sequence_len | 8192 | 512-32768 | Max token length. Longer = more VRAM (capped to 6144 on Intel XPU) |
micro_batch_size | 1 | limited by VRAM | Per-device batch size. H100: 7B QLoRA @8192 fits mb=8 |
gradient_accumulation_steps | 4 | 1-32 | Effective batch = micro * accum |
dataset_type | "alpaca" | "alpaca", "chat", "text" | |
val_set_size | 0.0 | 0-0.99 | Eval split; reports eval_loss/eval_perplexity |
lr_scheduler | "cosine" | "cosine", "constant" | With warmup_ratio (default 0.03) |
save_steps | null | ≥1 | Checkpoint cadence (resume-able jobs). null = auto (≈ total/20, min 200) |
flash_attention | "auto" | "auto", "on", "off" | Falls back to SDPA if not in image |
group_by_length | true | bool | Length-bucketed batches → less padding |
Full parameter reference (optimizer, gradient_checkpointing, sample_packing, pad_to_sequence_len, tokenizer_cache, seed, max_grad_norm, text_column…) in the LLM Fine-Tuning Guide.
VRAM optimization tips:
- On big cards, raise
micro_batch_sizebefore adding accumulation — it's the main throughput lever - Reduce
micro_batch_sizefirst (thensequence_len) if running out of VRAM qlorauses 4-bit quantization — much less VRAM thanloraornone
Time Series
TimesFM 2.5
Model ID: timesfm-2.5-finetune-gpu
| Parameter | Default | Range | Notes |
|---|---|---|---|
date_column | "ds" | — | Name of date/timestamp column |
target_column | "close" | — | Column to forecast. If absent, the first numeric column is used |
covariate_columns | [] | list of column names | Additional feature columns |
horizon_len | 128 | 1-1024 | Periods to forecast |
context_len | 1024 | 32-16384 | Historical periods to use as input |
epochs | 10 | 5-100 | Fine-tuning epochs |
batch_size | 8 | 1-64 | |
learning_rate | 1e-4 | 1e-6 to 1e-3 | |
quantiles | true | true, false | Quantile head (q10-q90) |
Prophet
Model ID: prophet-forecasting
| Parameter | Default | Options | Notes |
|---|---|---|---|
date_column | "ds" | — | Falls back to a column named ds, date, timestamp, time or datetime |
target_column | "y" | — | Column to forecast. If absent, the first numeric column is used |
horizon_len | 30 | 1-365 | Periods to forecast |
frequency | "D" | "D", "H", "W", "M", "Q", "T", "S" | Data frequency |
growth | "linear" | "linear", "logistic", "flat" | "logistic" needs cap |
cap | null | float | Carrying capacity for logistic growth |
floor | null | float | Lower saturation for logistic growth |
seasonality_mode | "additive" | "additive", "multiplicative" | Multiplicative for % growth patterns |
seasonality_prior_scale | 10.0 | 0.01-100 | Seasonality flexibility |
yearly_seasonality | "auto" | "auto", "true", "false" | Pass strings: a JSON boolean is treated as "auto" |
weekly_seasonality | "auto" | "auto", "true", "false" | Pass strings, as above |
daily_seasonality | "false" | "auto", "true", "false" | Only for sub-daily data |
country_holidays | "" | "US", "GB", "ES", etc. | Add country public holidays. Empty = none |
holidays_prior_scale | 10.0 | 0.01-100 | Holiday effect flexibility |
extra_regressors | [] | list of column names | Extra regressor columns |
changepoint_prior_scale | 0.05 | 0.001-0.5 | Trend flexibility. Higher = more flexible |
changepoint_range | 0.8 | 0.5-0.95 | Share of history where changepoints may occur |
n_changepoints | 25 | 0-100 | Number of potential changepoints |
interval_width | 0.8 | 0.5-0.99 | Width of the uncertainty interval |
mcmc_samples | 0 | 0-500 | 0 = MAP (fast). >0 = full Bayesian sampling |
ARIMA
Model ID: arima-forecasting
| Parameter | Default | Range | Notes |
|---|---|---|---|
date_column | null | — | null = auto-detect (first column that parses as dates) |
target_column | null | — | null = auto-detect (last numeric column) |
horizon | 30 | 1-365 | Periods to forecast |
frequency | "D" | "T", "H", "D", "W", "ME", "QE", "YE" | Data frequency |
seasonal | true | true, false | Enable SARIMA |
m | 1 | 1, 4, 7, 12, 52 | Seasonal period (12=monthly, 7=daily) |
information_criterion | "aic" | "aic", "bic", "hqic" | Model selection criterion |
max_p | 5 | 1-10 | Max AR order |
max_d | 2 | 0-3 | Max differencing order |
max_q | 5 | 1-10 | Max MA order |
Classical Forecasting (Auto)
Model ID: classical-forecasting-auto
| Parameter | Default | Notes |
|---|---|---|
date_column | "ds" | Falls back to a column named ds, date, timestamp, time or datetime |
target_columns | auto-detected | List of columns to forecast (or a single target_column, promoted to a list). Unset = one auto-detected target column |
horizon_len | 30 | Periods to forecast (1-365) |
context_len | 512 | History window length (10-2048) |
frequency | "D" | D, H, W, M, Q |
model_type | "auto" | "auto" selects Prophet, ARIMA or ETS per series; or force "prophet", "arima", "ets" |
Tuning Strategies
Start with Defaults
Always run with default hyperparameters first. For most tabular datasets, XGBoost defaults give strong results without tuning.
Use the Schema Endpoint
Get the exact parameter schema for any model. The path takes the template's UUID
(model_id), which GET /api/builder/v1/training/model-configs lists next to its name:
curl "https://api.colabhive.com/api/builder/v1/training/model-configs/$MODEL_ID/schema" \
-H "X-Account-ID: YOUR_ACCOUNT_ID" \
-H "X-API-Key: YOUR_API_KEY"
Key Levers (Tree-based)
-
Overfitting (train metrics great, validation poor):
- Reduce
max_depth(try 4 instead of 6) - Increase
reg_lambda(xgboost-regression,xgboost-classification) orlambda_l2(LightGBM) - Reduce
n_estimatorsor add early stopping - Reduce
num_leaves(LightGBM)
- Reduce
-
Underfitting (both train and validation metrics poor):
- Increase
n_estimators - Increase
max_depthornum_leaves(LightGBM) - Reduce
learning_rate+ increasen_estimators
- Increase
-
Slow training:
- Use GPU model (e.g.,
xgboost-regression-gpu) - Reduce
n_estimatorsfor quick experiments
- Use GPU model (e.g.,
See Also
- Training API — Model Config Schema — live parameter schema
- Preparing Datasets —
target_columnand format setup - LLM Fine-Tuning Guide — LoRA/QLoRA tuning detail
- Model Catalog — all available models