> For the complete documentation index, see [llms.txt](https://matterhorn-doc.mometic.com/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://matterhorn-doc.mometic.com/reference/methodology.md).

# Methodology

Understand Matterhorn deterministic scoring, point-in-time ML ranking, durability labels, missing data, and the limits of historical performance comparisons.

Matterhorn combines a traceable evidence system with two research views: existing deterministic scoring and an optional ML beta. Their common purpose is to prioritize investigation while retaining the basis for the result.

## Evidence before interpretation

Sources are stored with provenance, facts retain their source relationships, and calculations use explicit inputs. Validation controls check matters such as definitions, periods, units, consistency, and plausibility. These controls reduce errors; they do not guarantee that every accepted source interpretation is correct.

The ordinary company view evolves as the pipeline adds evidence. Historical ML research requires a stricter observation clock: a feature must have been publicly available by the company-date being modeled. Later filings, restatements, and outcomes must not be quietly inserted into an earlier view.

## The ML emergence target

The primary emergence ranker is **LightGBM LambdaRank**, grouped by historical observation month. The research target measures forward 24-month total return relative to the median observed return of the starting industry-and-size peer group, then ranks the adjusted outcomes across the month.

Industry classification and size must be supported at the starting observation. Missing or unresolved returns remain coverage gaps. The diagnostic winner/failure seed cohort is not the training label and does not represent the full training universe.

A high current percentile means the model prefers that input combination relative to others in the scored pool. It is neither a return forecast nor a calibrated probability of outperformance.

## The separate durability target

The initial definition starts from positive observed gross and operating margins. It tests whether operating income stays positive in each of the next four quarters and gross margin remains at least 90% of its starting level. A 40% starting gross margin implies a 36% floor under this definition.

The outcome is not mature until the required future filings are public. The model's output is conditional on its starting-economics population and is not a decades-long moat assessment. Where the target is inapplicable, the release formula uses its documented neutral handling rather than inventing persistence evidence.

## The fixed release blend

The release policy allocates 30% to the learned rank and 70% to the component aggregate. Within that component aggregate, the relative family weights are:

| Family              | Share of the component aggregate |
| ------------------- | -------------------------------: |
| Growth              |                          12 / 65 |
| Forward demand      |                          15 / 65 |
| Quality             |                          12 / 65 |
| Market confirmation |                           8 / 65 |
| Valuation           |                           6 / 65 |
| Evidence            |                           6 / 65 |
| Durability          |                           6 / 65 |

Multiply these proportions by 70% to obtain their contribution to the overall release formula. Family scores summarize supported feature percentiles; unavailable inputs retain neutral values under the fixed policy. The recipe is versioned. This blend is a release choice, not an empirically established universal optimum.

## Historical testing

Random train/test splitting is inappropriate for this task. The research design uses purged walk-forward evaluation, respects label maturity, compares against simpler baselines, and reserves a final holdout. Historical feature availability, delistings, ticker changes, corporate actions, and incomplete outcomes all affect the strength of the evidence.

The current beta should be treated as research software that has not graduated to automatic production ranking. A reproducible prediction and a passing software test demonstrate different things from out-of-sample predictive improvement. Investor-persona critiques can challenge assumptions but do not substitute for statistical validation or constitute endorsements from the named investors.

## Read performance claims with the right denominator

When comparing a top-10 selection process with SPY, require identical contribution dates, consistent holding rules, comparable dividend/corporate-action treatment, and explicit handling of delisted or unresolved names. Separate the number selected from the number with measured outcomes.

An average of overlapping two-year stock returns is not a continuously invested portfolio CAGR. With contributions on different dates, money-weighted return and ending wealth answer the investor's experience; time-weighted return describes the investment process without contribution timing. A common-date basket comparison is useful but does not, alone, prove a repeatable selection edge.

No historical return claim is promoted as a result of this user-guide release. The guide explains the live controls and how to evaluate their output.

Related: [ML enhanced](/understand-the-scores/ml-enhanced.md), [evidence quality](/understand-the-scores/evidence-quality.md), [glossary](/reference/glossary.md).
