Estimate whether cross-sectional return dispersion bends downward during larger market moves, with uncertainty and interpretation limits. The useful result is not just a scalar. You should be able to trace it back to eligible observations, explain which convention produced it and recognize when the calculation should stop.
This tutorial builds Equal-weight CSAD regression with intercept, absolute contemporaneous market return, squared market return, and Bartlett HAC errors. It is a descriptive herding diagnostic. You will calculate a small example, run matching Python and TypeScript implementations, inspect a controlled synthetic case and change one assumption in a guided lab. No historical market performance is claimed.
The chart shows the canonical fixture. Read the axis units before comparing values: a return, a score, a weight and a statistical diagnostic are different objects. Its numerical source is the same fixture used by the executable examples. Open the full-size chart when you need to inspect small labels.
Start with the question, then the mechanism
Without a common diagnostic plot, a negative coefficient can seem like a conclusion. The lab instead shows each market-return/dispersion pair and the fitted nonlinear curve. Compressing the cross-section in large-move periods bends the curve and changes its uncertainty. The construction is synthetic and deliberately known; actual evidence would still need to rule out common news, constraints, changing composition and misspecified market exposure.
Declare the sample, frequency, null hypothesis and tuning choices before inspecting test outcomes.
A diagnostic must name its null and its blind spots
Market efficiency is not one directly observable number. Different diagnostics study different restrictions: increment variance scaling, sign order, specified conditional-mean moments, event-relative returns or parameter stability. Rejecting one restriction does not establish a particular behavioral story. Failing to reject does not prove efficiency. Always state the null, sample, conditioning information, variant and alternative explanations beside the statistic.
Microstructure effects, asynchronous observations, stale prices, corporate actions and changing volatility can generate patterns that look like predictability. A diagnostic should therefore begin with a declared data frequency and return convention. A variance-ratio calculation uses additive increments, typically log returns; a compounded investment return uses simple returns. Converting between them is a modeling step, not a display-format change. Keep the distinction visible in the input contract.
Statistical choices have consequences. The runs test here uses an exact conditional distribution and a stated two-sided rule. The moment screen examines only two instruments and adjusts for those two tests. HAC-based regressions rely on asymptotic approximations and a declared lag count. The adaptive diagnostic uses a predeclared split; browsing many splits and reporting the smallest p-value would invalidate its nominal interpretation. A p-value is neither the probability that a theory is true nor the expected profit from a trade.
The behavioral topics are particularly careful about identification. A compressed cross-sectional dispersion curve does not observe imitation. A return-weighted expectations proxy does not observe stated beliefs. A post-event path does not identify overreaction without a credible benchmark, event definition and competing explanations. The labs let you isolate a mechanism with controlled synthetic interventions. Real empirical claims require a separately archived sample and an appropriate research design.
Freeze the definition
r is a common-universe return panel; the market proxy is the equal-weight cross-sectional mean; gamma_2 is the quadratic curvature coefficient.
The sources establish the method's research context; the stated variant fixes the implementation choices for this package. See Chang, Cheng and Khorana, An Examination of Herd Behavior, statsmodels, cov_hac documentation. Where a teaching convention differs from a published portfolio or test, it is labeled explicitly rather than borrowing the published method's empirical conclusions.
Work a small example before running the code
Security returns (-0.02,0.01,0.04) have equal-weight market return 0.01 and CSAD=(0.03+0+0.03)/3=0.02. A single date produces one dispersion point, not an estimate of herding. The regression needs many distinct market-move magnitudes.
The machine-readable hand check is saved separately from the larger chart fixture. It asserts coefficients against [0.0133333333, 0.666666667, 0]. Some hand checks use a different small input from the prose example to test the same invariant from another direction. For a model with several regressors, a one-row attribution example cannot estimate the loadings; the multi-period executable fixture supplies the necessary observations.
To audit the arithmetic, carry full precision through intermediate values and round only for display. Ask whether the result's unit is consistent with the formula. Then consider a limiting case: does the method return an explicit rejection or undefined result when its denominator or identifying variation disappears?
Prepare data without borrowing from the future
| Input | Type | Meaning |
|---|---|---|
| returns | number[][] | Dates by at least three securities, fixed universe and same return frequency. |
| hac_lags | integer | Bartlett covariance lag count, default 2. |
All calls also require formation_at, inputs_available_at and as_of as real ISO calendar dates. Inputs must be available by formation; formation cannot exceed the evaluation cutoff. Evaluation topics additionally require outcome start, end and availability dates. These envelope checks reject impossible chronology but cannot certify the provenance of individual rows. Your adapter must verify IDs, timestamps, frequency, currency, total-return adjustments, release dates and source vintages before building the arrays.
Missing, nonfinite, boolean or string-valued numbers are not silently repaired. The complete-case contract is intentional: changing eligibility changes the quantity being measured. Preserve the rejected records and the reason in a data-quality report, then choose a documented repair or a different model. Do not turn an undefined quantity into zero to make a chart look complete.
Follow the execution path
- Declare the information set. Declare the sample, frequency, null hypothesis and tuning choices before inspecting test outcomes.
- Build the diagnostic state. Retain the inputs and the intermediate quantities; alignment is part of correctness.
- Evaluate the null comparison. Calculate at full precision using the declared variant, not a convenient substitute.
- Inspect uncertainty and support. Check the method’s invariant and preserve undefined outcomes separately from numeric zero.
- Bound the conclusion. Statistical rejection applies to the declared null and sample. It does not identify investor motives, prove inefficiency or establish an executable profit.
Run the reference implementation
From the downloaded topic directory:
python examples/run.py
python -m unittest discover -s tests -p "test_*.py"
npx tsc -p implementations/typescript/tsconfig.json
node tests/test-typescript.mjs
The Python calculation has no third-party runtime dependency. TypeScript needs a compiler and an ES2022-capable JavaScript runtime. The public call accepts one JSON-shaped input and returns a discriminated success or error object. A minimal Python integration is:
from pathlib import Path
import importlib.util, json
root = Path.cwd() # Run from this topic directory.
spec = importlib.util.spec_from_file_location("topic", root / "implementations/python/algorithm.py")
module = importlib.util.module_from_spec(spec)
spec.loader.exec_module(module)
data = json.loads((root / "datasets/canonical-input.json").read_text())
result = module.compute(data)
if result["status"] != "ok":
raise ValueError(result["code"])
print(result["primary"])
After compilation, the equivalent TypeScript module can be used from JavaScript:
import { readFileSync } from 'node:fs';
import { compute } from './implementations/typescript/dist/algorithm.js';
const input = JSON.parse(readFileSync('./datasets/canonical-input.json', 'utf8'));
const result = compute(input);
if (result.status !== 'ok') throw new Error(result.code);
console.log(result.primary);
The canonical primary display is -4.69943059. More informative output fields include:
| Field | Canonical value / first values |
|---|---|
| coefficients | [0.010096242, 0.436547682, -4.69943059] |
| standard_errors | [0.000297142063, 0.0347031229, 0.764081169] |
| fitted | [0.010096242, 0.0154274981, 0.0184519345, 0.0197044638, 0.020028949, 0.0199339262, …] |
| residuals | [-9.62420215e-05, 0.000981754022, -0.000563923522, -0.000550782075, 0.000930596574, -9.85238108e-05, …] |
| r_squared | 0.946823738 |
Inspect the complete returned object rather than reducing every use case to primary. That field is a playground convenience; the named intermediate and result fields preserve the method's meaning. Both languages use the same defaults and reason codes and do not mutate the input. Package tests compare the whole output tree, while independent mathematical checks avoid treating one implementation as the sole authority for the other.
Use the playground as an experiment
Open the topic's Playground tab or the self-contained guided lab. It starts from a meaningful canonical preview. Choose a scenario, predict the result, use Step to follow the calculation, and explain the evidence before pressing Play. Back and Reset let you revisit exactly the same state. Reduced-motion mode advances one deliberate step instead of running a timed sequence.
The main control is Tail-compression coefficient, ranging from 0 to 8 with default 5. The comparison scenario is No quadratic compression. Downward fitted curvature describes dispersion shape; it does not establish imitation. Every change recomputes the result through the validated TypeScript kernel; it does not select a prerecorded result.
The deliberate failure scenario, No variation in market-move magnitude, should return SINGULAR_DESIGN. First explain which assumption failed. Then return to the canonical case and identify the information that makes the calculation possible. This rejection is part of the lesson: it prevents an invalid model from producing a plausible-looking number.
This second chart uses the comparison scenario at the default parameter. The caption and diagnostics in the lab explain what changes and what remains invariant. Identical output can be the correct outcome of an invariance experiment; do not mistake it for a broken control.
Avoid these interpretation failures
- A negative gamma_2 without uncertainty is not a detected phenomenon.
- Changing cross-sectional membership can change dispersion mechanically.
- The quadratic shape is only local to observed market moves; extrapolation can predict impossible negative dispersion.
Statistical rejection applies to the declared null and sample. It does not identify investor motives, prove inefficiency or establish an executable profit.
Check your understanding
Predict: Does a negative fitted quadratic coefficient identify investors copying each other?
Explain: No. It measures curvature conditional on this specification. Investor motives require additional identification and evidence.
Investigate: Run the canonical case, the comparison and the deliberate rejection. Save the input, output and one sentence explaining each difference. Identify a field whose unit could be confused with another field, and describe the consequence of that confusion.
Transfer: Before substituting real data, write the upstream eligibility and alignment rules. Name the source vintage, decision time and missing-value policy. Then identify one out-of-sample or data-quality check needed for your intended use. A successful synthetic calculation is a correctness demonstration, not evidence that the market rewards the signal.
What this package does and does not establish
The implementation makes the declared formula reproducible, exposes intermediates and rejects known invalid inputs. The sources motivate the method. The synthetic fixture lets you control one mechanism at a time. A named historical case remains deferred until its source observations and decision-time provenance can be archived; no invented returns are presented as real history.
Production use needs dataset-specific validation, monitored numerical limits, error logging, independent review and an execution or inference design appropriate to the application. See the source-package data contract and reference ledger for the full boundary. Educational material is not a recommendation to buy, sell or allocate capital.
Sources and further reading
- Chang, Cheng and Khorana, An Examination of Herd Behavior. CSAD nonlinear regression specification; a negative quadratic coefficient alone cannot establish imitation.
- statsmodels, cov_hac documentation. Bartlett-weight Newey-West covariance for equally spaced observations; our finite-sample convention is stated explicitly.
Choosing the method and continuing the lesson
This lesson is for analysts and developers who can work with aligned numerical arrays, means and return units. Regression and statistical-test topics also assume familiarity with residuals and sampling uncertainty; review the linked prerequisite before interpreting an inferential result.
| Decision | Declared approach | Neighbor or alternative |
|---|---|---|
| CSAD versus cross-sectional standard deviation | Averages absolute deviations. | Squares deviations and responds more strongly to extremes. |
| Negative curvature versus imitation | A fitted shape in return data. | A behavioral mechanism requiring evidence beyond this regression. |
Use the declared approach when its input and interpretation match your research question. If you choose the alternative, freeze a new convention and rerun the examples; changing a label is not enough to change the calculation.
Related concepts
Log return, Simple return. For any use with observed market data, keep the point-in-time dataset boundary explicit.
Learning connections
- Prerequisite: CAPM Beta. Establish the inputs or mathematical distinction used here.
- Comparison: Variance-Ratio Random-Walk Test. Compare its question and output units before substituting it for this method.
- Continue with: Overreaction and Underreaction Event Analysis. Carry the same formation clock and declared units into the next calculation.
Rendered from the canonical Mermaid sources linked by this article.
Calculation flow
ReferencesPrimary sources and evidence notesExpand the source trail, evidence role, and limitations behind the engineering choices.
Expand the source trail, evidence role, and limitations behind the engineering choices.
Chang, Cheng and Khorana, An Examination of Herd Behavior
- Source: Chang, Cheng and Khorana, An Examination of Herd Behavior
- Version / date: 2000
- Accessed: 2026-09-22
- Supports: CSAD nonlinear regression specification; a negative quadratic coefficient alone cannot establish imitation.
- Limitations: methodological context only; no claim that the source validates this synthetic sample or every educational convention.
- Reuse: cited, not copied. No source dataset is redistributed.
statsmodels, cov_hac documentation
- Source: statsmodels, cov_hac documentation
- Version / date: development documentation, accessed 2026-09-22
- Accessed: 2026-09-22
- Supports: Bartlett-weight Newey-West covariance for equally spaced observations; our finite-sample convention is stated explicitly.
- Limitations: methodological context only; no claim that the source validates this synthetic sample or every educational convention.
- Reuse: cited, not copied. No source dataset is redistributed.
Evidence boundaries
The formulas are operationalized in the canonical README with explicit package conventions. Original synthetic fixtures isolate mechanisms and are not a historical performance claim. External source access can be restricted; the MacKinlay archive is a bibliographic reference, not a claim that its full text was retrieved during this build.
Historical case decision: deferred. A named empirical case would require a separately archived point-in-time universe, source vintage and outcome design. A synthetic control is used here to demonstrate dispersion curvature without attributing invented observations to a market. This limits empirical coverage; it does not change the mathematical contract.
When using live data, archive the retrieval date, provider query, license, currency, frequency, adjustment basis and transformation log. Do not imply that the primary authors endorsed this educational implementation.
Full dependency-light reference implementations in both supported languages.
/** D17 reference algorithms. JSON boundary validation is deliberate and shared.
* Arrays are copied before sorting; callers' inputs are never mutated.
* QR solves least squares without forming normal equations.
*/
type Data = Record<string, any>;
type Result = Record<string, any>;
class ContractError extends Error {
}
const fail = (code: string): never => { throw new ContractError(code); };
const num = (x: unknown): number => typeof x === 'number' && Number.isFinite(x) ? x : fail('INVALID_NUMBER');
function integer(x: unknown, lo: number, hi: number): number { const v = num(x); return Number.isInteger(v) && v >= lo && v <= hi ? v : fail('INVALID_PARAMETER'); }
function vec(x: unknown, min = 1): number[] { if (!Array.isArray(x))
fail('INVALID_SHAPE'); const a = x as unknown[]; if (a.length < min)
fail('INSUFFICIENT_DATA'); return a.map(num); }
function mat(x: unknown, min = 1): number[][] { if (!Array.isArray(x) || x.length < min)
fail('INSUFFICIENT_DATA'); const a = (x as unknown[]).map(v => vec(v)); if (new Set(a.map(r => r.length)).size !== 1)
fail('LENGTH_MISMATCH'); return a; }
function same(...x: {
length: number;
}[]): void { if (new Set(x.map(a => a.length)).size !== 1)
fail('LENGTH_MISMATCH'); }
function ids(d: Data, n: number): string[] { required(d, ['ids']); if (!Array.isArray(d.ids) || d.ids.length !== n)
fail('LENGTH_MISMATCH'); if (d.ids.some((v: unknown) => typeof v !== 'string' || !/^[A-Za-z0-9_.-]+$/.test(v)))
fail('INVALID_ID'); if (new Set(d.ids).size !== n)
fail('DUPLICATE_ID'); return [...d.ids]; }
function required(d: Data, keys: string[]): void { for (const key of keys)
if (!Object.hasOwn(d, key))
fail('MISSING_FIELD'); }
function validDate(v: unknown): boolean { if (typeof v !== 'string' || !/^\d{4}-\d{2}-\d{2}$/.test(v) || v.startsWith('0000'))
return false; const date = new Date(v + 'T00:00:00Z'); return Number.isFinite(date.valueOf()) && date.toISOString().slice(0, 10) === v; }
function context(d: Data, op: string): void {
const keys = ['as_of', 'formation_at', 'inputs_available_at'];
const evaluation = ['ic', 'rank_ic', 'spread', 'decay'].includes(op);
if (evaluation)
keys.push('outcomes_start_at', 'outcomes_end_at', 'outcomes_available_at');
required(d, keys);
if (keys.some(k => !validDate(d[k])))
fail('INVALID_DATE');
if (d.formation_at > d.as_of || d.inputs_available_at > d.formation_at)
fail('FUTURE_INPUT');
if (evaluation) {
if (d.outcomes_start_at < d.formation_at || d.outcomes_end_at < d.outcomes_start_at)
fail('INVALID_OUTCOME_WINDOW');
if (d.outcomes_available_at < d.outcomes_end_at)
fail('INVALID_DATE_ORDER');
if (d.outcomes_available_at > d.as_of)
fail('IMMATURE_OUTCOME');
}
}
const sum = (a: number[]): number => a.reduce((s, v) => s + v, 0);
const mean = (a: number[]): number => sum(a) / a.length;
const dot = (a: number[], b: number[]): number => sum(a.map((v, i) => v * b[i]));
const tr = (a: number[][]): number[][] => a[0].map((_, j) => a.map(r => r[j]));
const mm = (a: number[][], b: number[][]): number[][] => { const cols = tr(b); return a.map(row => cols.map(col => dot(row, col))); };
function sd(x: number[], ddof = 0): number { const m = mean(x); return Math.sqrt(sum(x.map(v => (v - m) ** 2)) / (x.length - ddof)); }
function standard(x: number[]): number[] { const m = mean(x), s = sd(x); if (s <= 1e-14 * Math.max(1, ...x.map(Math.abs)))
fail('CONSTANT_CROSS_SECTION'); return x.map(v => (v - m) / s); }
function corr(x: number[], y: number[]): number | null { same(x, y); if (Math.max(...x) === Math.min(...x) || Math.max(...y) === Math.min(...y))
return null; const xm = mean(x), ym = mean(y), a = x.map(v => v - xm), b = y.map(v => v - ym), den = Math.sqrt(dot(a, a) * dot(b, b)); return den <= 0 ? null : Math.max(-1, Math.min(1, dot(a, b) / den)); }
function ranks(x: number[]): number[] { const order = x.map((_, i) => i).sort((a, b) => x[a] - x[b]); const out = x.map(() => 0); let start = 0; while (start < x.length) {
let end = start + 1;
while (end < x.length && x[order[end]] === x[order[start]])
end++;
for (let j = start; j < end; j++)
out[order[j]] = (start + 1 + end) / 2;
start = end;
} return out; }
function quantile(x: number[], p: number): number { const y = [...x].sort((a, b) => a - b), h = (y.length - 1) * p, j = Math.floor(h), f = h - j; return y[j] * (1 - f) + y[Math.min(j + 1, y.length - 1)] * f; }
function compound(x: number[]): number { if (x.some(v => v <= -1))
fail('INVALID_RETURN'); return Math.expm1(sum(x.map(Math.log1p))); }
/** erfc(|z|/sqrt(2)) via regularized Gamma(1/2,x); converged series/CF. */
function normalP(z: number): number {
const x = z * z / 2, a = 0.5, lg = 0.5723649429247001;
if (x === 0)
return 1;
const factor = Math.exp(-x + a * Math.log(x) - lg);
if (x < a + 1) {
let term = 1 / a, total = term, ap = a;
for (let i = 1; i < 500; i++) {
ap++;
term *= x / ap;
total += term;
if (Math.abs(term) < Math.abs(total) * 1e-15)
break;
}
return Math.max(0, 1 - total * factor);
}
let b = x + 1 - a, c = 1e300, d = 1 / b, h = d;
for (let i = 1; i < 500; i++) {
const an = -i * (i - a);
b += 2;
d = an * d + b;
if (Math.abs(d) < 1e-300)
d = 1e-300;
c = b + an / c;
if (Math.abs(c) < 1e-300)
c = 1e-300;
d = 1 / d;
const delta = d * c;
h *= delta;
if (Math.abs(delta - 1) < 1e-15)
break;
}
return Math.max(0, Math.min(1, factor * h));
}
export function ols(y: number[], x: number[][], lags = 0): Result {
const n = y.length, p = x[0].length;
same(y, x);
if (n <= p)
fail('INSUFFICIENT_DATA');
integer(lags, 0, n - 1);
const cols = tr(x), scales = cols.map(c => Math.sqrt(dot(c, c)));
if (scales.some(s => s === 0))
fail('SINGULAR_DESIGN');
const q: number[][] = [], r = Array.from({ length: p }, () => Array(p).fill(0) as number[]);
for (let j = 0; j < p; j++) {
let v = cols[j].map(z => z / scales[j]);
for (let pass = 0; pass < 2; pass++)
for (let i = 0; i < j; i++) {
const proj = dot(q[i], v);
r[i][j] += proj;
v = v.map((z, t) => z - proj * q[i][t]);
}
r[j][j] = Math.sqrt(dot(v, v));
if (r[j][j] < 1e-10)
fail('SINGULAR_DESIGN');
q.push(v.map(z => z / r[j][j]));
}
const solve = (v: number[]): number[] => { const b = Array(p).fill(0) as number[]; for (let i = p - 1; i >= 0; i--) {
let s = 0;
for (let j = i + 1; j < p; j++)
s += r[i][j] * b[j];
b[i] = (v[i] - s) / r[i][i];
} return b; };
const beta = solve(q.map(c => dot(c, y))).map((b, i) => b / scales[i]), fitted = x.map(row => dot(row, beta)), residuals = y.map((v, i) => v - fitted[i]);
const invr = tr(Array.from({ length: p }, (_, j) => solve(Array.from({ length: p }, (_, i) => Number(i === j)))));
const bread = mm(invr, tr(invr)).map((row, i) => row.map((v, j) => v / scales[i] / scales[j]));
const scores = x.map((row, t) => row.map(v => v * residuals[t])), meat = mm(tr(scores), scores);
for (let lag = 1; lag <= lags; lag++) {
const w = 1 - lag / (lags + 1);
for (let t = lag; t < n; t++)
for (let i = 0; i < p; i++)
for (let j = 0; j < p; j++)
meat[i][j] += w * (scores[t][i] * scores[t - lag][j] + scores[t - lag][i] * scores[t][j]);
}
const cov = mm(mm(bread, meat), bread), se = cov.map((row, i) => Math.sqrt(Math.max(0, row[i]))), sse = dot(residuals, residuals), ym = mean(y), sst = sum(y.map(v => (v - ym) ** 2));
return { coefficients: beta, standard_errors: se, fitted, residuals, r_squared: sst === 0 ? null : 1 - sse / sst, n, df_residual: n - p, hac_lags: lags, residual_sum_squares: sse, qr_min_diagonal: Math.min(...r.map((row, i) => row[i])) };
}
function hacVariance(g: number[], lags: number): number { const n = g.length, m = mean(g), a = g.map(v => v - m); let v = dot(a, a) / n; for (let lag = 1; lag <= lags; lag++)
v += 2 * (1 - lag / (lags + 1)) * dot(a.slice(lag), a.slice(0, -lag)) / n; return v; }
function choose(n: number, k: number): number { if (k < 0 || k > n)
return 0; k = Math.min(k, n - k); let out = 1; for (let j = 1; j <= k; j++)
out *= (n - k + j) / j; return out; }
function behavior(d: Data, op: string): Result {
if (op === 'herding') {
required(d, ['returns']);
const panel = mat(d.returns, 12);
if (panel.some(row => row.some(v => v < -1)))
fail('INVALID_RETURN');
if (panel[0].length < 3)
fail('INSUFFICIENT_DATA');
const market = panel.map(mean), csad = panel.map((row, t) => mean(row.map(v => Math.abs(v - market[t])))), fit = ols(csad, market.map(m => [1, Math.abs(m), m * m]), d.hac_lags ?? 2), se = fit.standard_errors[2], z = se <= 1e-14 ? null : fit.coefficients[2] / se;
return { ...fit, market, csad, curvature: fit.coefficients[2], z, p_value: z === null ? null : normalP(z), primary: fit.coefficients[2] };
}
if (op === 'event') {
required(d, ['estimation_asset', 'estimation_market', 'event_asset', 'event_market', 'event_times', 'estimation_end', 'event_start']);
const a = vec(d.estimation_asset, 8), m = vec(d.estimation_market, 8), ea = vec(d.event_asset), em = vec(d.event_market), times = vec(d.event_times);
same(a, m);
same(ea, em, times);
if (Math.min(...a, ...m, ...ea, ...em) < -1)
fail('INVALID_RETURN');
if (!validDate(d.estimation_end) || !validDate(d.event_start))
fail('INVALID_DATE');
if (d.estimation_end >= d.event_start)
fail('OVERLAPPING_WINDOWS');
if (d.event_start > d.inputs_available_at)
fail('FUTURE_INPUT');
if (times.some((t, i) => !Number.isInteger(t) || (i > 0 && t <= times[i - 1])))
fail('INVALID_ORDER');
const fit = ols(a, m.map(v => [1, v]), 0), expected = em.map(v => fit.coefficients[0] + fit.coefficients[1] * v), ar = ea.map((v, i) => v - expected[i]);
let total = 0;
const car = ar.map(v => (total += v));
return { coefficients: fit.coefficients, expected_returns: expected, abnormal_returns: ar, car, event_times: times, primary: car[car.length - 1] };
}
required(d, ['returns']);
const r = vec(d.returns, 2), n = r.length;
if (['expectations', 'adaptive'].includes(op) && Math.min(...r) < -1)
fail('INVALID_RETURN');
if (op === 'variance_ratio') {
const q = integer(d.q ?? 2, 2, n - 1), mu = mean(r), e = r.map(v => v - mu), ss = dot(e, e);
if (Math.max(...r) === Math.min(...r) || ss <= 0)
fail('CONSTANT_SERIES');
const blocks = Array.from({ length: n - q + 1 }, (_, j) => sum(r.slice(j, j + q)) - q * mu), short = ss / (n - 1), long = dot(blocks, blocks) / (q * (n - q + 1) * (1 - q / n)), vr = long / short, deltas = Array.from({ length: q - 1 }, (_, k) => { const j = k + 1; let v = 0; for (let t = j; t < n; t++)
v += e[t] ** 2 * e[t - j] ** 2; return v / ss ** 2; }), theta = sum(deltas.map((v, k) => 4 * (1 - (k + 1) / q) ** 2 * v));
if (theta <= 0)
fail('ZERO_MOMENT_VARIANCE');
const z = (vr - 1) / Math.sqrt(theta);
return { ratio: vr, z, p_value: normalP(z), q, one_step_variance: short, q_step_variance: long, robust_variance: theta, deltas, block_sums: blocks, primary: vr };
}
if (op === 'runs') {
const threshold = num(d.threshold ?? 0), signs = r.filter(v => v !== threshold).map(v => v > threshold ? 1 : -1), a = signs.filter(v => v === 1).length, b = signs.length - a, total = a + b;
if (!a || !b)
fail('ONE_SIGN');
if (total > 200)
fail('EXACT_SAMPLE_LIMIT');
const runs = 1 + sum(signs.slice(1).map((v, i) => Number(v !== signs[i]))), expected = 1 + 2 * a * b / total, variance = 2 * a * b * (2 * a * b - total) / (total ** 2 * (total - 1)), den = choose(total, a), probs: Array<{
runs: number;
probability: number;
}> = [];
for (let k = 2; k <= total; k++) {
let ways: number;
if (k % 2 === 0) {
const j = k / 2;
ways = 2 * choose(a - 1, j - 1) * choose(b - 1, j - 1);
}
else {
const j = (k - 1) / 2;
ways = choose(a - 1, j) * choose(b - 1, j - 1) + choose(a - 1, j - 1) * choose(b - 1, j);
}
probs.push({ runs: k, probability: ways / den });
}
const observed = probs.find(p => p.runs === runs)!.probability, p = Math.min(1, sum(probs.filter(row => row.probability <= observed * (1 + 1e-12)).map(row => row.probability)));
return { runs, expected, variance, z: variance === 0 ? null : (runs - expected) / Math.sqrt(variance), p_value: p, signs, zeros_dropped: n - total, distribution: probs, primary: p };
}
if (op === 'martingale') {
if (n < 12)
fail('INSUFFICIENT_DATA');
const lags = integer(d.hac_lags ?? 2, 0, n - 2), moments = [1, 2].map(power => r.slice(1).map((v, t) => v * r[t] ** power)), omega = moments.map(g => hacVariance(g, lags));
if (Math.min(...omega) <= 0)
fail('ZERO_MOMENT_VARIANCE');
const z = moments.map((g, i) => Math.sqrt(n - 1) * mean(g) / Math.sqrt(omega[i])), p = z.map(normalP), family = Math.min(1, 2 * Math.min(...p));
return { moments, means: moments.map(mean), long_run_variances: omega, z, p_values: p, family_p_value: family, hac_lags: lags, primary: family };
}
if (op === 'horizons') {
const a = integer(d.recent ?? 1, 1, n), b = integer(d.middle ?? 12, a + 1, n), c = integer(d.distant ?? 36, b + 1, n);
if (Math.min(...r) <= -1)
fail('INVALID_RETURN');
const blocks = [r.slice(n - a), r.slice(n - b, n - a), r.slice(n - c, n - b)], logs = blocks.map(row => sum(row.map(Math.log1p))), total = Math.expm1(sum(logs));
return { log_contributions: logs, simple_returns: logs.map(Math.expm1), candidate_signals: [-logs[0], logs[1], -logs[2]], total_return: total, endpoints: [a, b, c], primary: total };
}
if (op === 'expectations') {
const lookback = integer(d.lookback ?? 12, 1, n), rho = num(d.rho ?? 0.8);
if (rho < 0 || rho > 1)
fail('INVALID_PARAMETER');
const observed = r.slice(-lookback).reverse(), raw = Array.from({ length: lookback }, (_, j) => rho ** j), weights = raw.map(v => v / sum(raw)), contributions = weights.map((v, i) => v * observed[i]);
return { weights, returns_recent_first: observed, contributions, proxy: sum(contributions), primary: sum(contributions), rho };
}
if (op === 'adaptive') {
if (n < 24)
fail('INSUFFICIENT_DATA');
const split = integer(d.split ?? Math.floor(n / 2), 9, n - 8), design = r.slice(1).map((_, j) => { const t = j + 1, D = Number(t >= split); return [1, r[t - 1], D, D * r[t - 1]]; }), fit = ols(r.slice(1), design, d.hac_lags ?? 2), se = fit.standard_errors[3], z = se <= 1e-14 ? null : fit.coefficients[3] / se;
return { ...fit, pre_beta: fit.coefficients[1], post_beta: fit.coefficients[1] + fit.coefficients[3], change: fit.coefficients[3], z, p_value: z === null ? null : normalP(z), split, primary: fit.coefficients[3] };
}
return fail('UNKNOWN_METHOD');
}
/** D17-F05-A05 public boundary. No mutation, implicit imputation or silent failure. */
export function compute(input: unknown): Result {
const op = "herding";
try {
if (!input || typeof input !== 'object' || Array.isArray(input))
fail('INVALID_SHAPE');
const d = input as Data;
if (Object.values(d).some(v => v === null))
fail('INVALID_NUMBER');
context(d, op);
const result = behavior(d, op);
const check = (v: unknown): void => { if (typeof v === 'number' && !Number.isFinite(v))
fail('NUMERICAL_FAILURE'); if (Array.isArray(v))
v.forEach(check);
else if (v && typeof v === 'object')
Object.values(v).forEach(check); };
check(result);
return { status: 'ok', method: op, ...result };
}
catch (error) {
if (error instanceof ContractError)
return { status: 'error', method: op, code: error.message };
throw error;
}
}
The embedded lab now expands to its full document height, keeping the article as the only scroll surface.
