Skip to main content

Hyperparameter Tuning

Reference for all configurable hyperparameters per model family, with defaults, ranges, and tuning strategies.

Tip: Use GET /training/model-configs/{model_id}/schema to get the live JSON Schema for any model's hyperparameters — always up to date.

Target column (tabular templates): target_column is not a declared schema parameter. The training node promotes it to the declared list parameter target_columns ("price" → ["price"]), so it only takes effect on templates that declare target_columns; each table says which. With no target given, the target is auto-detected: the analyzed dataset schema's target, else a column named target, label, output or y, else the last column.

Seeds: the tabular and BERT templates that declare seed default to null, which keeps each template's built-in behavior. Pass an integer for reproducible runs.


Tree-Based Models (XGBoost, LightGBM, CatBoost, Random Forest)​

XGBoost Regression / Classification​

Model IDs: xgboost-regression, xgboost-regression-gpu, xgboost-classification

ParameterDefaultRangeNotes
n_estimators10050-5000xgboost-regression, xgboost-classification only. Number of trees. More = better but slower
n_estimators10010-1000xgboost-regression-gpu only. Number of trees (this ID accepts at most 1000)
max_depth63-12Tree depth. Higher = more complex patterns, more overfitting risk
learning_rate0.10.01-0.5Step size. Lower LR + more trees = better generalization
subsample1.00.5-1.0Fraction of samples per tree. Values below 1.0 reduce overfitting
colsample_bytree1.00.5-1.0Fraction of features per tree
reg_lambda10-10xgboost-regression, xgboost-classification only. L2 regularization
objectivenull"binary:logistic", "multi:softprob"xgboost-classification only. Learning objective. null picks binary:logistic for a two-class target and multi:softprob otherwise. Just these two probability objectives are accepted, because the template's evaluation and the serving path read class probabilities. binary:logistic needs a two-class target; multi:softprob returns one probability per class
eval_metricnull"auc", "aucpr", "error", "logloss", "merror", "mlogloss"xgboost-classification only. Metric XGBoost logs on the training data every boosting round. The template does not stop early, so the metric does not change the fitted model. null means logloss with binary:logistic and mlogloss with multi:softprob. error and logloss need binary:logistic, merror and mlogloss need multi:softprob, and auc and aucpr work with both. A mismatch with an explicit objective is refused before queuing; with objective unset, it fails before training
min_child_weight1.00 or morexgboost-classification only. Minimum sum of instance weight (hessian) needed in a child. Higher = more conservative trees
scale_pos_weightnullgreater than 0xgboost-classification only. Weight of positive relative to negative examples for binary:logistic, for unbalanced classes; a typical value is the number of negatives divided by the number of positives. null keeps XGBoost's default, 1. It must be greater than 0 (0 itself is refused). Refused with multi:softprob, and a run whose target has more than two classes fails before training
seednull0-2147483647Random seed. null keeps the built-in behavior
target_columnauto-detected—Name of the column to predict. Promoted to target_columns
target_columns—list of column namesxgboost-regression, xgboost-regression-gpu only. Several target columns (multi-output)
target_columnsnullone column name, as a one-item listxgboost-classification only. One explicit target column; null auto-detects it. target_column is its singular alias, and conflicting values are refused before queuing

Quick presets:

# Small dataset (<10k) — more regularization
hyperparameters={
"n_estimators": 200,
"max_depth": 4,
"learning_rate": 0.05,
"subsample": 0.8,
}

# Large dataset (>100k) — more trees
hyperparameters={
"n_estimators": 1000,
"max_depth": 6,
"learning_rate": 0.05,
"subsample": 0.8,
}

LightGBM Regression / Classification​

Model IDs: lightgbm-regression-gpu, lightgbm-classification

ParameterDefaultRangeNotes
n_estimators100100-5000Number of boosting rounds
num_leaves3120-255Max leaves per tree (key LightGBM param)
max_depth61-20lightgbm-regression-gpu only. Depth cap applied on top of num_leaves
max_depth-1-1 (unlimited) to 50lightgbm-classification only. -1 means controlled by num_leaves
learning_rate0.10.01-0.3
feature_fraction1.00.1-1.0Fraction of features per tree (LightGBM's name for colsample_bytree)
lambda_l20.00-10L2 regularization
subsample1.00.5-1.0lightgbm-classification only. Fraction of samples per tree (bagging turns on below 1.0)
seednull0-2147483647Random seed. null keeps the built-in behavior
target_columnauto-detected—Promoted to target_columns
target_columns—list of column nameslightgbm-regression-gpu only. Several target columns (multi-output)
target_columnsnullone column name, as a one-item listlightgbm-classification only. One explicit target column; null auto-detects it. target_column is its singular alias, and conflicting values are refused before queuing

Key difference vs XGBoost: Use num_leaves instead of max_depth. Start from the default num_leaves=31 and tune from there.


Random Forest Regression / Classification​

Model IDs: random-forest-regression, random-forest-classification

ParameterDefaultRangeNotes
n_estimators10050-500Number of trees
max_depth101-50Max tree depth. Higher = deeper, more specific trees
min_samples_split22-20Min samples to split a node
min_samples_leaf11-10Min samples at leaf
seednull0-2147483647Random seed. null keeps the built-in behavior
target_columnauto-detected—Promoted to target_columns
target_columns—list of column namesrandom-forest-regression only. Several target columns (multi-output)
target_columnsnullone column name, as a one-item listrandom-forest-classification only. One explicit target column; null auto-detects it. target_column is its singular alias, and conflicting values are refused before queuing

CatBoost Regression / Classification​

Model IDs: catboost-regression, catboost-classification

ParameterDefaultRangeNotes
n_estimators500100-5000Number of trees (CatBoost's iterations)
max_depth64-10Tree depth (CatBoost's depth)
learning_rate0.030.01-0.3
l2_leaf_reg31-10L2 regularization
subsample1.00.1-1.0catboost-regression only. Fraction of samples per tree
subsamplenullgreater than 0, up to 1catboost-classification only. Bagging sample rate; it must be greater than 0 (0 itself is refused) and at most 1. null keeps CatBoost's default bootstrap: MVS at 0.8 for a binary target on CPU with 100 rows or more, and Bayesian, which has no sample rate, for a multi-class target and on GPU. When set, a binary target on CPU keeps MVS; elsewhere the template switches to Bernoulli bootstrap. effective_parameters.json records the bootstrap type and rate used
seednull0-2147483647Random seed. null keeps the built-in behavior
target_columnauto-detected—Promoted to target_columns
target_columns—list of column namescatboost-regression only. Target column list
target_columnsnullone column name, as a one-item listcatboost-classification only. One explicit target column; null auto-detects it. target_column is its singular alias, and conflicting values are refused before queuing

CatBoost handles categorical features natively — no need to encode string columns before training.


Linear & Logistic Regression​

Model IDs: linear-regression, logistic-regression

The two templates declare different parameters.

Linear Regression​

Model ID: linear-regression

ParameterDefaultOptionsNotes
regularization"ridge""ridge", "lasso", "elasticnet", "none"Type of regularization
alpha1.00.001-100Regularization strength. Higher = stronger
l1_ratio0.50-1ElasticNet mix (0=Ridge, 1=Lasso)
fit_intercepttruetrue, falseFit an intercept term
target_columnauto-detected—Ignored by this template: the target is auto-detected

Logistic Regression​

Model ID: logistic-regression

ParameterDefaultOptionsNotes
C1.00.0001-1000Inverse regularization strength. Lower = stronger regularization
penalty"l2""l1", "l2", "elasticnet", "none"Type of regularization
l1_ratio0.50-1ElasticNet mix, used when penalty is "elasticnet"
max_iter1000100-10000Max solver iterations
fit_intercepttruetrue, falseFit an intercept term
seednull0-2147483647Random seed. null keeps the built-in behavior
target_columnauto-detected—Ignored by this template: the target is auto-detected

SVM / SVR​

Model IDs: svm-classification, svr-regression

ParameterDefaultOptionsNotes
kernel"rbf""rbf", "linear", "poly", "sigmoid"Kernel type
C1.00.01-1000Regularization. Higher = less regularized
gamma"scale""scale", "auto"Kernel coefficient (numeric values are not accepted)
degree32-10Polynomial degree, used by the "poly" kernel
epsilon0.10-10svr-regression only. Width of the tube with no penalty
seednull0-2147483647svm-classification only. Random seed. null keeps the built-in behavior
target_columnauto-detected—Ignored by these templates: the target is auto-detected

SVM is slow on large datasets (>50k rows). Prefer XGBoost or LightGBM for larger data.


Deep Learning (MLP, TabNet)​

MLP (Multi-Layer Perceptron)​

Model ID: mlp-tabular-gpu

ParameterDefaultRangeNotes
hidden_dims[128, 64, 32]List of intsLayer sizes
learning_rate0.0010.0001-0.01Adam LR
batch_size6416-512Training batch size
epochs5010-500Training epochs
target_columnauto-detected—Promoted to target_columns
target_columns—list of column namesSeveral target columns (multi-output)

TabNet​

Model ID: tabnet-tabular-gpu

ParameterDefaultRangeNotes
n_steps31-10Number of sequential attention steps
n_d84-64Decision embedding dimension
n_a84-64Attention embedding dimension
gamma1.31.0-2.0Feature reuse across steps. Higher = more reuse allowed
learning_rate0.020.001-0.1
batch_size25664-1024
epochs10050-500
target_columnauto-detected—Promoted to target_columns
target_columns—list of column namesSeveral target columns (multi-output)

NLP (BERT Classification)​

Model ID: bert-classification-gpu

ParameterDefaultRangeNotes
epochs32-20Fine-tuning epochs
batch_size168-128Reduce if OOM errors
learning_rate2e-51e-6 to 5e-5Typical BERT range
max_length12864-512Max token length per sample
text_column"text"—Column with text data
label_column"label"—Column with class labels
base_model"bert-base-uncased"—Hugging Face encoder to fine-tune
seednull0-2147483647Random seed. null keeps the built-in behavior

LLM Fine-Tuning (llm-qlora-finetune)​

Model ID: llm-qlora-finetune

See the LLM Fine-Tuning Guide for a full walkthrough.

ParameterDefaultOptions / RangeNotes
base_modelQwen/Qwen2.5-Coder-7B-InstructAny compatible HF CausalLMHuggingFace repo ID
adapter"qlora""qlora", "lora", "none"none = full fine-tune (needs more VRAM)
epochsnull1-10null = 3 epochs (alias num_epochs)
learning_rate0.00021e-5 to 5e-4Typical LoRA range
lora_r164-128LoRA rank. Higher = more parameters
lora_alpha328-256Usually lora_r * 2
lora_dropout0.050-0.2
sequence_len8192512-32768Max token length. Longer = more VRAM (capped to 6144 on Intel XPU)
micro_batch_size1limited by VRAMPer-device batch size. H100: 7B QLoRA @8192 fits mb=8
gradient_accumulation_steps41-32Effective batch = micro * accum
dataset_type"alpaca""alpaca", "chat", "text"
val_set_size0.00-0.99Eval split; reports eval_loss/eval_perplexity
lr_scheduler"cosine""cosine", "constant"With warmup_ratio (default 0.03)
save_stepsnull≥1Checkpoint cadence (resume-able jobs). null = auto (≈ total/20, min 200)
flash_attention"auto""auto", "on", "off"Falls back to SDPA if not in image
group_by_lengthtrueboolLength-bucketed batches → less padding

Full parameter reference (optimizer, gradient_checkpointing, sample_packing, pad_to_sequence_len, tokenizer_cache, seed, max_grad_norm, text_column…) in the LLM Fine-Tuning Guide.

VRAM optimization tips:

  • On big cards, raise micro_batch_size before adding accumulation — it's the main throughput lever
  • Reduce micro_batch_size first (then sequence_len) if running out of VRAM
  • qlora uses 4-bit quantization — much less VRAM than lora or none

Time Series​

TimesFM 2.5​

Model ID: timesfm-2.5-finetune-gpu

ParameterDefaultRangeNotes
date_column"ds"—Name of date/timestamp column
target_column"close"—Column to forecast. If absent, the first numeric column is used
covariate_columns[]list of column namesAdditional feature columns
horizon_len1281-1024Periods to forecast
context_len102432-16384Historical periods to use as input
epochs105-100Fine-tuning epochs
batch_size81-64
learning_rate1e-41e-6 to 1e-3
quantilestruetrue, falseQuantile head (q10-q90)

Prophet​

Model ID: prophet-forecasting

ParameterDefaultOptionsNotes
date_column"ds"—Falls back to a column named ds, date, timestamp, time or datetime
target_column"y"—Column to forecast. If absent, the first numeric column is used
horizon_len301-365Periods to forecast
frequency"D""D", "H", "W", "M", "Q", "T", "S"Data frequency
growth"linear""linear", "logistic", "flat""logistic" needs cap
capnullfloatCarrying capacity for logistic growth
floornullfloatLower saturation for logistic growth
seasonality_mode"additive""additive", "multiplicative"Multiplicative for % growth patterns
seasonality_prior_scale10.00.01-100Seasonality flexibility
yearly_seasonality"auto""auto", "true", "false"Pass strings: a JSON boolean is treated as "auto"
weekly_seasonality"auto""auto", "true", "false"Pass strings, as above
daily_seasonality"false""auto", "true", "false"Only for sub-daily data
country_holidays"""US", "GB", "ES", etc.Add country public holidays. Empty = none
holidays_prior_scale10.00.01-100Holiday effect flexibility
extra_regressors[]list of column namesExtra regressor columns
changepoint_prior_scale0.050.001-0.5Trend flexibility. Higher = more flexible
changepoint_range0.80.5-0.95Share of history where changepoints may occur
n_changepoints250-100Number of potential changepoints
interval_width0.80.5-0.99Width of the uncertainty interval
mcmc_samples00-5000 = MAP (fast). >0 = full Bayesian sampling

ARIMA​

Model ID: arima-forecasting

ParameterDefaultRangeNotes
date_columnnull—null = auto-detect (first column that parses as dates)
target_columnnull—null = auto-detect (last numeric column)
horizon301-365Periods to forecast
frequency"D""T", "H", "D", "W", "ME", "QE", "YE"Data frequency
seasonaltruetrue, falseEnable SARIMA
m11, 4, 7, 12, 52Seasonal period (12=monthly, 7=daily)
information_criterion"aic""aic", "bic", "hqic"Model selection criterion
max_p51-10Max AR order
max_d20-3Max differencing order
max_q51-10Max MA order

Classical Forecasting (Auto)​

Model ID: classical-forecasting-auto

ParameterDefaultNotes
date_column"ds"Falls back to a column named ds, date, timestamp, time or datetime
target_columnsauto-detectedList of columns to forecast (or a single target_column, promoted to a list). Unset = one auto-detected target column
horizon_len30Periods to forecast (1-365)
context_len512History window length (10-2048)
frequency"D"D, H, W, M, Q
model_type"auto""auto" selects Prophet, ARIMA or ETS per series; or force "prophet", "arima", "ets"

Tuning Strategies​

Start with Defaults​

Always run with default hyperparameters first. For most tabular datasets, XGBoost defaults give strong results without tuning.

Use the Schema Endpoint​

Get the exact parameter schema for any model. The path takes the template's UUID (model_id), which GET /api/builder/v1/training/model-configs lists next to its name:

curl "https://api.colabhive.com/api/builder/v1/training/model-configs/$MODEL_ID/schema" \
-H "X-Account-ID: YOUR_ACCOUNT_ID" \
-H "X-API-Key: YOUR_API_KEY"

Key Levers (Tree-based)​

  1. Overfitting (train metrics great, validation poor):

    • Reduce max_depth (try 4 instead of 6)
    • Increase reg_lambda (xgboost-regression, xgboost-classification) or lambda_l2 (LightGBM)
    • Reduce n_estimators or add early stopping
    • Reduce num_leaves (LightGBM)
  2. Underfitting (both train and validation metrics poor):

    • Increase n_estimators
    • Increase max_depth or num_leaves (LightGBM)
    • Reduce learning_rate + increase n_estimators
  3. Slow training:

    • Use GPU model (e.g., xgboost-regression-gpu)
    • Reduce n_estimators for quick experiments

See Also​