Skip to main content

Principal Component Analysis (PCA)

TYPE 'pca' projects the feature vector onto the top-k principal components computed via Jacobi eigendecomposition of the centred covariance matrix. Training accumulates covariance statistics, then an iterative Jacobi eigensolver finds the component basis. Use PCA to compress a wide feature set before clustering / regression / nearest-neighbour search, or for exploratory visualization. The native projection centers the data but does not whiten it.

When to use it​

  • Feature sets where many columns are correlated and a smaller orthogonal basis preserves most of the variance.
  • Pre-processing before KMeans — KMeans assumes isotropic clusters and benefits from decorrelated inputs.
  • Building dense vector representations from high-dimensional sparse-ish inputs.

When NOT to use it​

  • Categorical features without one-hot encoding plus careful scaling (PCA assumes meaningful Euclidean structure).
  • Very wide feature sets: the covariance matrix alone needs quadratic memory in the feature count. Consider external dimensionality reduction when that cost is prohibitive.
  • The downstream task is non-linear and a tree-based model would handle the original features fine — PCA is unnecessary.

Syntax​

CREATE MODEL <name>(DOUBLE, DOUBLE[, ...]) RETURNS ARRAY<DOUBLE>
TYPE { 'pca' | 'principal_component_analysis' }
OPTIONS (...)
AS SELECT <f1>, <f2>, ... FROM <source>;

PCA is unsupervised — every SELECT column is treated as a feature, no label.

Options​

OptionDefaultTypeRangeWhat it does
num_components (alias k)min(dim, 8)int1 <= k <= dimNumber of principal components to keep
max_jacobi_sweeps30int>= 1Maximum sweeps in the Jacobi eigensolver
jacobi_tolerance1e-10float> 0Off-diagonal-norm convergence threshold for the eigensolver
partitions1int>= 1Training partitions; see execution and memory limits

Examples​

These fragments assume the named source tables exist. Match training and inference feature order and preprocessing. For a complete dataset and runnable script, follow the linked tutorial.

Minimal — compress 5 features into 2:

CREATE MODEL features_2d(DOUBLE, DOUBLE, DOUBLE, DOUBLE, DOUBLE) RETURNS ARRAY<DOUBLE>
TYPE 'pca'
OPTIONS (num_components = 2)
AS SELECT f1, f2, f3, f4, f5 FROM raw_features;

SELECT id, features_2d(f1, f2, f3, f4, f5) AS pc FROM raw_features;

Tuned with tighter eigensolver convergence and distributed training:

CREATE MODEL embedding_compressor(DOUBLE, DOUBLE, DOUBLE, DOUBLE, DOUBLE, DOUBLE, DOUBLE, DOUBLE) RETURNS ARRAY<DOUBLE>
TYPE 'pca'
OPTIONS (
k = 4,
max_jacobi_sweeps = 100,
jacobi_tolerance = 1e-12,
partitions = 4
)
AS SELECT e1, e2, e3, e4, e5, e6, e7, e8
FROM document_embeddings;

Output shape​

For two or more components, one model call returns ARRAY<DOUBLE> with one element per component. For exactly one component, declare RETURNS DOUBLE. If both num_components and k are supplied, they must agree.

WITH projected AS (
SELECT id, features_2d(f1, f2, f3, f4, f5) AS pcs FROM raw_features
)
SELECT id, pcs[1] AS pc1, pcs[2] AS pc2 FROM projected;

Tuning notes​

  • Standardise features (z-score) before PCA — otherwise the largest-variance column dominates the basis.
  • Pick k from a scree plot: the elbow where the explained-variance ratio flattens.
  • max_jacobi_sweeps = 30 and jacobi_tolerance = 1e-10 are fine for most workloads; tighten only if downstream metrics demand it.
  • For very wide inputs the closed-form Jacobi step gets expensive — bucket the input or pre-aggregate before training.
  • partitions > 1 parallelises the covariance accumulation but the eigensolve still runs once on the merged matrix.

Convergence and quality​

The outer training summary does not establish projection quality. Check output length, finite values, variation along retained axes, and whether downstream performance remains acceptable. Component signs can flip between equivalent fits; compare projections with that ambiguity in mind. A small-variance direction may still contain useful target information, so retained variance alone is not a supervised quality metric.