# Methodology — how we detect changes

- **TrackLLM** — machine-readable twin of [methodology.html](https://www.trackllm.net/methodology.html). Every page on the site has a `.md` sibling at the same path; links below point at those.
- **Raw data** (every collected response): https://github.com/timothee-chauvin/trackllm_data · Built 2026-10-05 09:12 UTC

[Home](index.md) / Methodology

TrackLLM uses two methods, depending on what the API gives us back.

## LT — Logprob tracking

Used wherever the provider returns logprobs. Paper: [Log Probability Tracking of LLM APIs](https://arxiv.org/abs/2512.03816) (ICLR 2026 · arXiv:2512.03816).

When the API returns logprobs, they make for extremely sensitive and cost-effective change detection. LT is a simple statistical test on the average value of each token's logprob, requesting only a single token of output.

Logprobs are noisy, because inference on GPUs is non-deterministic, but their averages over a few dozen queries are stable enough to track. The test detects changes as small as one step of fine-tuning, and is more sensitive than existing methods while being 1,000x cheaper — logprobs can be used to monitor for changes at extremely low cost, e.g. $0.14/year for hourly sampling of GPT-4.1.

## B3IT — Black-box border input tracking

Used where logprobs aren't exposed. Paper: [Token-Efficient Change Detection in LLM APIs](https://arxiv.org/abs/2602.11083) (ICML 2026 · arXiv:2602.11083).

Where logprobs aren't exposed, we operate in a strict black box, observing only output tokens. The first phase identifies border inputs: inputs for which sampling at T=0 doesn't always give the same output, i.e. for which there exists more than one output top token. They can easily be found just from black-box sampling, trying thousands of short inputs and keeping the border inputs. These inputs are then sampled repeatedly, to check whether the endpoint has moved away from its borders.

Optimal change detection depends on the model's Jacobian and the Fisher information of the output distribution; analyzing these in low-temperature regimes shows that border inputs enable powerful change detection tests. B3IT performs on par with the best available gray-box approaches while reducing costs by 30x.

## Read more

- [Change detection in LLM APIs](https://tchauvin.com/change-detection-llm-apis) — blog post, the short version of both methods.
- [Log Probability Tracking of LLM APIs](https://arxiv.org/abs/2512.03816) — ICLR 2026 · arXiv:2512.03816 — the LT test, and the TinyChange benchmark.
- [Token-Efficient Change Detection in LLM APIs](https://arxiv.org/abs/2602.11083) — ICML 2026 · arXiv:2602.11083 — border inputs and B3IT.

---
Data collected via the OpenRouter API. [Home](https://www.trackllm.net/index.md) · [Changes](https://www.trackllm.net/changes.md) · [Providers](https://www.trackllm.net/providers.md) · [Endpoints](https://www.trackllm.net/endpoints.md) · [Methodology](https://www.trackllm.net/methodology.md) · [About](https://www.trackllm.net/about.md) · [GitHub](https://www.trackllm.net/github.md)
