Rank lower historical total-return variability while showing the lookback and annualization that define it. The useful result is not just a scalar. You should be able to trace it back to eligible observations, explain which convention produced it and recognize when the calculation should stop.
This tutorial builds Negative realized sample volatility standardized across securities. This is a standalone characteristic, not minimum-variance portfolio optimization. You will calculate a small example, run matching Python and TypeScript implementations, inspect a controlled synthetic case and change one assumption in a guided lab. No historical market performance is claimed.
The chart shows the canonical fixture. Read the axis units before comparing values: a return, a score, a weight and a statistical diagnostic are different objects. Its numerical source is the same fixture used by the executable examples. Open the full-size chart when you need to inspect small labels.
Start with the question, then the mechanism
Volatility depends on deviations around a mean, not on whether the average return was high or low. Two stocks can earn the same cumulative return with very different paths. The sample denominator L−1 is explicit because short windows make denominator choices visible. After calculating volatility, the negative sign aligns the score so that calmer histories rank higher. The lab changes the lookback to reveal a shock that enters or leaves the measurement window.
Use only information available at formation; standardize within the explicitly supplied eligible cross section using population standard deviation.
A characteristic becomes a score only after a convention is chosen
A cross-sectional score is an ordering device for a formation universe. It does not measure an expected return in percentage points. Its value depends on which entities are eligible, what information was available, and how the characteristic is transformed. A company can receive a different standardized score with unchanged fundamentals when its peers change. Save the universe and raw characteristic next to the score so that the difference is explainable.
Accounting period end is not information availability. A December balance sheet published in March must not enter a January formation. Preserve the filing release timestamp, reporting currency, units, consolidation basis and restatement vintage in the upstream adapter. The compact teaching API checks a declared availability envelope; it does not verify the provenance of each accounting row. That responsibility remains with the data pipeline. Price-based histories need their own calendar, corporate-action and total-return policy.
These examples use complete, finite inputs and reject invalid denominators. They never silently replace missing characteristics with zero. If a research design allows negative book equity or missing leverage, define a separate variant and its eligibility rule before seeing forward returns. Likewise, a winsorization policy changes the population being described; it must be an explicit upstream construction decision rather than an undocumented rescue inside the score function.
Population standardization subtracts the cross-sectional mean and divides by dispersion using N. A zero-dispersion characteristic carries no relative ordering information and is rejected by z-score methods. A large score is a relative deviation, not a confidence level. After formation, use the construction family to choose portfolios and the evaluation family to examine matured outcomes. This separation keeps future returns out of the score definition.
Freeze the definition
L is at least two observations; A is the declared observations per year; v is annualized sample volatility in decimal return units.
The sources establish the method's research context; the stated variant fixes the implementation choices for this package. See MSCI, Foundations of Factor Investing. Where a teaching convention differs from a published portfolio or test, it is labeled explicitly rather than borrowing the published method's empirical conclusions.
Work a small example before running the code
Returns (-0.01,0.01) have sample variance 0.0002. With A=12, annualized volatility is sqrt(0.0024)=0.04898979. A larger A rescales all volatilities equally, so cross-sectional z-scores do not change when all firms use the same frequency.
The machine-readable hand check is saved separately from the larger chart fixture. It asserts raw against [-0.0489897949, -0.0979795897, -0.146969385]. Some hand checks use a different small input from the prose example to test the same invariant from another direction. For a model with several regressors, a one-row attribution example cannot estimate the loadings; the multi-period executable fixture supplies the necessary observations.
To audit the arithmetic, carry full precision through intermediate values and round only for display. Ask whether the result's unit is consistent with the formula. Then consider a limiting case: does the method return an explicit rejection or undefined result when its denominator or identifying variation disappears?
Prepare data without borrowing from the future
| Input | Type | Meaning |
|---|---|---|
| ids | string[] | Unique security identifiers. |
| returns | number[][] | Aligned oldest-to-newest total-return histories. |
| lookback | integer | Number of final observations to use. |
| annualization | number | Positive periods per year; monthly default 12. |
All calls also require formation_at, inputs_available_at and as_of as real ISO calendar dates. Inputs must be available by formation; formation cannot exceed the evaluation cutoff. Evaluation topics additionally require outcome start, end and availability dates. These envelope checks reject impossible chronology but cannot certify the provenance of individual rows. Your adapter must verify IDs, timestamps, frequency, currency, total-return adjustments, release dates and source vintages before building the arrays.
Missing, nonfinite, boolean or string-valued numbers are not silently repaired. The complete-case contract is intentional: changing eligibility changes the quantity being measured. Preserve the rejected records and the reason in a data-quality report, then choose a documented repair or a different model. Do not turn an undefined quantity into zero to make a chart look complete.
Follow the execution path
- Freeze eligible names. Use only information available at formation; standardize within the explicitly supplied eligible cross section using population standard deviation.
- Calculate the raw characteristic. Retain the inputs and the intermediate quantities; alignment is part of correctness.
- Check denominator and coverage. Calculate at full precision using the declared variant, not a convenient substitute.
- Standardize across the universe. Check the method’s invariant and preserve undefined outcomes separately from numeric zero.
- Compare score and portfolio meaning. A score is a characteristic used in ranking or portfolio design; it is not a realized factor return or evidence of a premium.
Run the reference implementation
From the downloaded topic directory:
python examples/run.py
python -m unittest discover -s tests -p "test_*.py"
npx tsc -p implementations/typescript/tsconfig.json
node tests/test-typescript.mjs
The Python calculation has no third-party runtime dependency. TypeScript needs a compiler and an ES2022-capable JavaScript runtime. The public call accepts one JSON-shaped input and returns a discriminated success or error object. A minimal Python integration is:
from pathlib import Path
import importlib.util, json
root = Path.cwd() # Run from this topic directory.
spec = importlib.util.spec_from_file_location("topic", root / "implementations/python/algorithm.py")
module = importlib.util.module_from_spec(spec)
spec.loader.exec_module(module)
data = json.loads((root / "datasets/canonical-input.json").read_text())
result = module.compute(data)
if result["status"] != "ok":
raise ValueError(result["code"])
print(result["primary"])
After compilation, the equivalent TypeScript module can be used from JavaScript:
import { readFileSync } from 'node:fs';
import { compute } from './implementations/typescript/dist/algorithm.js';
const input = JSON.parse(readFileSync('./datasets/canonical-input.json', 'utf8'));
const result = compute(input);
if (result.status !== 'ok') throw new Error(result.code);
console.log(result.primary);
The canonical primary display is -0.523722856. More informative output fields include:
| Field | Canonical value / first values |
|---|---|
| ids | ["SYN-001", "SYN-002", "SYN-003", "SYN-004", "SYN-005", "SYN-006", …] |
| raw | [-0.0557446064, -0.0561570134, -0.0420156422, -0.0358233872, -0.0497750763, -0.0596493929, …] |
| scores | [-0.523722856, -0.609514814, 2.33227813, 3.62043693, 0.718103077, -1.33602549, …] |
| window_start | 24 |
| window_end_exclusive | 36 |
Inspect the complete returned object rather than reducing every use case to primary. That field is a playground convenience; the named intermediate and result fields preserve the method's meaning. Both languages use the same defaults and reason codes and do not mutate the input. Package tests compare the whole output tree, while independent mathematical checks avoid treating one implementation as the sole authority for the other.
Use the playground as an experiment
Open the topic's Playground tab or the self-contained guided lab. It starts from a meaningful canonical preview. Choose a scenario, predict the result, use Step to follow the calculation, and explain the evidence before pressing Play. Back and Reset let you revisit exactly the same state. Reduced-motion mode advances one deliberate step instead of running a timed sequence.
The main control is Included monthly observations, ranging from 2 to 30 with default 12. The comparison scenario is Most recent shock. A shock changes the chosen window's total volatility; this is not market beta. Every change recomputes the result through the validated TypeScript kernel; it does not select a prerecorded result.
The deliberate failure scenario, Simple return below minus 100 percent, should return INVALID_RETURN. First explain which assumption failed. Then return to the canonical case and identify the information that makes the calculation possible. This rejection is part of the lesson: it prevents an invalid model from producing a plausible-looking number.
This second chart uses the comparison scenario at the default parameter. The caption and diagnostics in the lab explain what changes and what remains invariant. Identical output can be the correct outcome of an invariance experiment; do not mistake it for a broken control.
Avoid these interpretation failures
- Mixing daily and monthly histories defeats a common annualization factor.
- Stale prices can create deceptively low measured volatility.
- A quiet sample is not a guarantee of low future risk.
A score is a characteristic used in ranking or portfolio design; it is not a realized factor return or evidence of a premium.
Check your understanding
Predict: Is the lowest total-volatility name necessarily the lowest-beta name?
Explain: No. Total volatility includes residual variation; beta depends on covariance with a specified market and its variance.
Investigate: Run the canonical case, the comparison and the deliberate rejection. Save the input, output and one sentence explaining each difference. Identify a field whose unit could be confused with another field, and describe the consequence of that confusion.
Transfer: Before substituting real data, write the upstream eligibility and alignment rules. Name the source vintage, decision time and missing-value policy. Then identify one out-of-sample or data-quality check needed for your intended use. A successful synthetic calculation is a correctness demonstration, not evidence that the market rewards the signal.
What this package does and does not establish
The implementation makes the declared formula reproducible, exposes intermediates and rejects known invalid inputs. The sources motivate the method. The synthetic fixture lets you control one mechanism at a time. A named historical case remains deferred until its source observations and decision-time provenance can be archived; no invented returns are presented as real history.
Production use needs dataset-specific validation, monitored numerical limits, error logging, independent review and an execution or inference design appropriate to the application. See the source-package data contract and reference ledger for the full boundary. Educational material is not a recommendation to buy, sell or allocate capital.
Sources and further reading
- MSCI, Foundations of Factor Investing. Construction choices matter; this package does not reproduce an MSCI index or its current methodology.
Choosing the method and continuing the lesson
This lesson is for analysts and developers who can work with aligned numerical arrays, means and return units. Regression and statistical-test topics also assume familiarity with residuals and sampling uncertainty; review the linked prerequisite before interpreting an inferential result.
| Decision | Declared approach | Neighbor or alternative |
|---|---|---|
| Low volatility versus low beta | Measures total historical variability. | Measures sensitivity to a selected market return. |
| Low-vol score versus minimum variance | Orders individual securities. | Optimizes portfolio risk using correlations and constraints. |
Use the declared approach when its input and interpretation match your research question. If you choose the alternative, freeze a new convention and rerun the examples; changing a label is not enough to change the calculation.
Related concepts
Factor score, Standardized score. For any use with observed market data, keep the point-in-time dataset boundary explicit.
Learning connections
- Preparation: review the linked glossary definitions, means, dispersion and the input contract before starting.
- Comparison: CAPM Beta. Compare its question and output units before substituting it for this method.
- Continue with: Profitability Factor Score. Carry the same formation clock and declared units into the next calculation.
Rendered from the canonical Mermaid sources linked by this article.
Calculation flow
ReferencesPrimary sources and evidence notesExpand the source trail, evidence role, and limitations behind the engineering choices.
Expand the source trail, evidence role, and limitations behind the engineering choices.
MSCI, Foundations of Factor Investing
- Source: MSCI, Foundations of Factor Investing
- Version / date: 2013
- Accessed: 2026-09-22
- Supports: Construction choices matter; this package does not reproduce an MSCI index or its current methodology.
- Limitations: methodological context only; no claim that the source validates this synthetic sample or every educational convention.
- Reuse: cited, not copied. No source dataset is redistributed.
Evidence boundaries
The formulas are operationalized in the canonical README with explicit package conventions. Original synthetic fixtures isolate mechanisms and are not a historical performance claim. External source access can be restricted; the MacKinlay archive is a bibliographic reference, not a claim that its full text was retrieved during this build.
Historical case decision: deferred. A named empirical case would require a separately archived point-in-time universe, source vintage and outcome design. A synthetic control is used here to demonstrate cross section without attributing invented observations to a market. This limits empirical coverage; it does not change the mathematical contract.
When using live data, archive the retrieval date, provider query, license, currency, frequency, adjustment basis and transformation log. Do not imply that the primary authors endorsed this educational implementation.
Full dependency-light reference implementations in both supported languages.
/** D17 reference algorithms. JSON boundary validation is deliberate and shared.
* Arrays are copied before sorting; callers' inputs are never mutated.
* QR solves least squares without forming normal equations.
*/
type Data = Record<string, any>;
type Result = Record<string, any>;
class ContractError extends Error {
}
const fail = (code: string): never => { throw new ContractError(code); };
const num = (x: unknown): number => typeof x === 'number' && Number.isFinite(x) ? x : fail('INVALID_NUMBER');
function integer(x: unknown, lo: number, hi: number): number { const v = num(x); return Number.isInteger(v) && v >= lo && v <= hi ? v : fail('INVALID_PARAMETER'); }
function vec(x: unknown, min = 1): number[] { if (!Array.isArray(x))
fail('INVALID_SHAPE'); const a = x as unknown[]; if (a.length < min)
fail('INSUFFICIENT_DATA'); return a.map(num); }
function mat(x: unknown, min = 1): number[][] { if (!Array.isArray(x) || x.length < min)
fail('INSUFFICIENT_DATA'); const a = (x as unknown[]).map(v => vec(v)); if (new Set(a.map(r => r.length)).size !== 1)
fail('LENGTH_MISMATCH'); return a; }
function same(...x: {
length: number;
}[]): void { if (new Set(x.map(a => a.length)).size !== 1)
fail('LENGTH_MISMATCH'); }
function ids(d: Data, n: number): string[] { required(d, ['ids']); if (!Array.isArray(d.ids) || d.ids.length !== n)
fail('LENGTH_MISMATCH'); if (d.ids.some((v: unknown) => typeof v !== 'string' || !/^[A-Za-z0-9_.-]+$/.test(v)))
fail('INVALID_ID'); if (new Set(d.ids).size !== n)
fail('DUPLICATE_ID'); return [...d.ids]; }
function required(d: Data, keys: string[]): void { for (const key of keys)
if (!Object.hasOwn(d, key))
fail('MISSING_FIELD'); }
function validDate(v: unknown): boolean { if (typeof v !== 'string' || !/^\d{4}-\d{2}-\d{2}$/.test(v) || v.startsWith('0000'))
return false; const date = new Date(v + 'T00:00:00Z'); return Number.isFinite(date.valueOf()) && date.toISOString().slice(0, 10) === v; }
function context(d: Data, op: string): void {
const keys = ['as_of', 'formation_at', 'inputs_available_at'];
const evaluation = ['ic', 'rank_ic', 'spread', 'decay'].includes(op);
if (evaluation)
keys.push('outcomes_start_at', 'outcomes_end_at', 'outcomes_available_at');
required(d, keys);
if (keys.some(k => !validDate(d[k])))
fail('INVALID_DATE');
if (d.formation_at > d.as_of || d.inputs_available_at > d.formation_at)
fail('FUTURE_INPUT');
if (evaluation) {
if (d.outcomes_start_at < d.formation_at || d.outcomes_end_at < d.outcomes_start_at)
fail('INVALID_OUTCOME_WINDOW');
if (d.outcomes_available_at < d.outcomes_end_at)
fail('INVALID_DATE_ORDER');
if (d.outcomes_available_at > d.as_of)
fail('IMMATURE_OUTCOME');
}
}
const sum = (a: number[]): number => a.reduce((s, v) => s + v, 0);
const mean = (a: number[]): number => sum(a) / a.length;
const dot = (a: number[], b: number[]): number => sum(a.map((v, i) => v * b[i]));
const tr = (a: number[][]): number[][] => a[0].map((_, j) => a.map(r => r[j]));
const mm = (a: number[][], b: number[][]): number[][] => { const cols = tr(b); return a.map(row => cols.map(col => dot(row, col))); };
function sd(x: number[], ddof = 0): number { const m = mean(x); return Math.sqrt(sum(x.map(v => (v - m) ** 2)) / (x.length - ddof)); }
function standard(x: number[]): number[] { const m = mean(x), s = sd(x); if (s <= 1e-14 * Math.max(1, ...x.map(Math.abs)))
fail('CONSTANT_CROSS_SECTION'); return x.map(v => (v - m) / s); }
function corr(x: number[], y: number[]): number | null { same(x, y); if (Math.max(...x) === Math.min(...x) || Math.max(...y) === Math.min(...y))
return null; const xm = mean(x), ym = mean(y), a = x.map(v => v - xm), b = y.map(v => v - ym), den = Math.sqrt(dot(a, a) * dot(b, b)); return den <= 0 ? null : Math.max(-1, Math.min(1, dot(a, b) / den)); }
function ranks(x: number[]): number[] { const order = x.map((_, i) => i).sort((a, b) => x[a] - x[b]); const out = x.map(() => 0); let start = 0; while (start < x.length) {
let end = start + 1;
while (end < x.length && x[order[end]] === x[order[start]])
end++;
for (let j = start; j < end; j++)
out[order[j]] = (start + 1 + end) / 2;
start = end;
} return out; }
function quantile(x: number[], p: number): number { const y = [...x].sort((a, b) => a - b), h = (y.length - 1) * p, j = Math.floor(h), f = h - j; return y[j] * (1 - f) + y[Math.min(j + 1, y.length - 1)] * f; }
function compound(x: number[]): number { if (x.some(v => v <= -1))
fail('INVALID_RETURN'); return Math.expm1(sum(x.map(Math.log1p))); }
/** erfc(|z|/sqrt(2)) via regularized Gamma(1/2,x); converged series/CF. */
function normalP(z: number): number {
const x = z * z / 2, a = 0.5, lg = 0.5723649429247001;
if (x === 0)
return 1;
const factor = Math.exp(-x + a * Math.log(x) - lg);
if (x < a + 1) {
let term = 1 / a, total = term, ap = a;
for (let i = 1; i < 500; i++) {
ap++;
term *= x / ap;
total += term;
if (Math.abs(term) < Math.abs(total) * 1e-15)
break;
}
return Math.max(0, 1 - total * factor);
}
let b = x + 1 - a, c = 1e300, d = 1 / b, h = d;
for (let i = 1; i < 500; i++) {
const an = -i * (i - a);
b += 2;
d = an * d + b;
if (Math.abs(d) < 1e-300)
d = 1e-300;
c = b + an / c;
if (Math.abs(c) < 1e-300)
c = 1e-300;
d = 1 / d;
const delta = d * c;
h *= delta;
if (Math.abs(delta - 1) < 1e-15)
break;
}
return Math.max(0, Math.min(1, factor * h));
}
export function ols(y: number[], x: number[][], lags = 0): Result {
const n = y.length, p = x[0].length;
same(y, x);
if (n <= p)
fail('INSUFFICIENT_DATA');
integer(lags, 0, n - 1);
const cols = tr(x), scales = cols.map(c => Math.sqrt(dot(c, c)));
if (scales.some(s => s === 0))
fail('SINGULAR_DESIGN');
const q: number[][] = [], r = Array.from({ length: p }, () => Array(p).fill(0) as number[]);
for (let j = 0; j < p; j++) {
let v = cols[j].map(z => z / scales[j]);
for (let pass = 0; pass < 2; pass++)
for (let i = 0; i < j; i++) {
const proj = dot(q[i], v);
r[i][j] += proj;
v = v.map((z, t) => z - proj * q[i][t]);
}
r[j][j] = Math.sqrt(dot(v, v));
if (r[j][j] < 1e-10)
fail('SINGULAR_DESIGN');
q.push(v.map(z => z / r[j][j]));
}
const solve = (v: number[]): number[] => { const b = Array(p).fill(0) as number[]; for (let i = p - 1; i >= 0; i--) {
let s = 0;
for (let j = i + 1; j < p; j++)
s += r[i][j] * b[j];
b[i] = (v[i] - s) / r[i][i];
} return b; };
const beta = solve(q.map(c => dot(c, y))).map((b, i) => b / scales[i]), fitted = x.map(row => dot(row, beta)), residuals = y.map((v, i) => v - fitted[i]);
const invr = tr(Array.from({ length: p }, (_, j) => solve(Array.from({ length: p }, (_, i) => Number(i === j)))));
const bread = mm(invr, tr(invr)).map((row, i) => row.map((v, j) => v / scales[i] / scales[j]));
const scores = x.map((row, t) => row.map(v => v * residuals[t])), meat = mm(tr(scores), scores);
for (let lag = 1; lag <= lags; lag++) {
const w = 1 - lag / (lags + 1);
for (let t = lag; t < n; t++)
for (let i = 0; i < p; i++)
for (let j = 0; j < p; j++)
meat[i][j] += w * (scores[t][i] * scores[t - lag][j] + scores[t - lag][i] * scores[t][j]);
}
const cov = mm(mm(bread, meat), bread), se = cov.map((row, i) => Math.sqrt(Math.max(0, row[i]))), sse = dot(residuals, residuals), ym = mean(y), sst = sum(y.map(v => (v - ym) ** 2));
return { coefficients: beta, standard_errors: se, fitted, residuals, r_squared: sst === 0 ? null : 1 - sse / sst, n, df_residual: n - p, hac_lags: lags, residual_sum_squares: sse, qr_min_diagonal: Math.min(...r.map((row, i) => row[i])) };
}
function score(d: Data, op: string): Result {
let raw: number[];
if (op === 'value') {
required(d, ['book_equity', 'market_equity']);
const b = vec(d.book_equity, 3), m = vec(d.market_equity, 3);
same(b, m);
if (Math.min(...b, ...m) <= 0)
fail('INVALID_DENOMINATOR');
raw = b.map((v, i) => v / m[i]);
}
else if (op === 'size') {
required(d, ['market_equity']);
const m = vec(d.market_equity, 3);
if (Math.min(...m) <= 0)
fail('INVALID_DENOMINATOR');
raw = m.map(v => -Math.log(v));
}
else if (op === 'profitability') {
required(d, ['revenue', 'cogs', 'assets']);
const a = vec(d.revenue, 3), b = vec(d.cogs, 3), c = vec(d.assets, 3);
same(a, b, c);
if (Math.min(...c) <= 0 || Math.min(...a, ...b) < 0)
fail('INVALID_DENOMINATOR');
raw = a.map((v, i) => (v - b[i]) / c[i]);
}
else if (op === 'quality') {
required(d, ['roe', 'debt_equity', 'earnings_variability']);
const a = vec(d.roe, 3), b = vec(d.debt_equity, 3), c = vec(d.earnings_variability, 3);
same(a, b, c);
if (Math.min(...b, ...c) < 0)
fail('INVALID_DENOMINATOR');
const components = tr([standard(a), standard(b).map(v => -v), standard(c).map(v => -v)]);
raw = components.map(mean);
return { ids: ids(d, raw.length), raw, scores: raw, components, primary: raw[0] };
}
else {
required(d, ['returns']);
const histories = mat(d.returns, 3), lookback = integer(d.lookback ?? (op === 'momentum' ? 11 : 12), op === 'momentum' ? 1 : 2, histories[0].length), skip = op === 'momentum' ? integer(d.skip ?? 1, 0, histories[0].length - lookback) : 0;
if (histories.some(row => row.some(v => v < -1 || (op === 'momentum' && v === -1))))
fail('INVALID_RETURN');
const end = histories[0].length - skip, included = histories.map(row => row.slice(end - lookback, end));
if (op === 'momentum')
raw = included.map(compound);
else {
const annual = num(d.annualization ?? 12);
if (annual <= 0)
fail('INVALID_PARAMETER');
raw = included.map(row => -sd(row, 1) * Math.sqrt(annual));
}
const scores = standard(raw);
return { ids: ids(d, raw.length), raw, scores, primary: scores[0], window_start: end - lookback, window_end_exclusive: end, lookback, skip };
}
const scores = standard(raw);
return { ids: ids(d, raw.length), raw, scores, primary: scores[0] };
}
/** D17-F02-A05 public boundary. No mutation, implicit imputation or silent failure. */
export function compute(input: unknown): Result {
const op = "lowvol";
try {
if (!input || typeof input !== 'object' || Array.isArray(input))
fail('INVALID_SHAPE');
const d = input as Data;
if (Object.values(d).some(v => v === null))
fail('INVALID_NUMBER');
context(d, op);
const result = score(d, op);
const check = (v: unknown): void => { if (typeof v === 'number' && !Number.isFinite(v))
fail('NUMERICAL_FAILURE'); if (Array.isArray(v))
v.forEach(check);
else if (v && typeof v === 'object')
Object.values(v).forEach(check); };
check(result);
return { status: 'ok', method: op, ...result };
}
catch (error) {
if (error instanceof ContractError)
return { status: 'error', method: op, code: error.message };
throw error;
}
}
The embedded lab now expands to its full document height, keeping the article as the only scroll surface.
