Results#
The sealed board#
PROTEA’s headline result on the leakage-free temporal frame is a weighted,
IA-aware micro F-measure (f_micro_w) of 0.40765, first in seven of the
nine evaluation cells (NK/LK/PK by BPO/MFO/CCO).
Frame. Test window v227 to v230; validation window v225 to v227. The window is temporal by construction: the reference pool is frozen at t0 and the ground truth is the delta gained by t1, so the score is free of the data leakage that inflates tools scored against a current database.
Champion. The learned k-WTA retrieval encoder (config
d8979601, which stores GO-aligned codes rather than a raw PLM vector, a choice the representation ablation in Operational Insights and Lessons Learned motivates) for candidate generation, followed by a stacked per-category re-ranker (see ADR-D43: Stacked meta-reranker (evidence scorers plus a shallow per-category combiner)).Metric and scoring.
f_micro_wis the headline metric, scored board-faithfully withcafaevalon the sealed settings (OBOreleases/2025-07-22, the t0 IA artefact, the release terms-of-interest,prop=fill,norm=cafa,no_orphans, and-knownon PK only so Partial-Knowledge excludes already-known terms). The protocol, the per-band OBO/IA registry, and the LAFA parity mapping are documented in full in CAFA Evaluation Protocol, which is the home of the metric definition and the scoring recipe.
The two cells not won are LK-BPO and PK-BPO: the Biological Process wall. It is a
limit of the ranking stage, not of the available evidence. 97.0 percent of the true
terms missed on those cells already exist in the pre-window vocabulary, 95.2 percent
by the information accretion the metric weights by, and the candidate pool the
pipeline already retrieves is worth f_micro_w 0.7519 at precision 1.000 to a
perfect ranker, and up to 0.7764 to the best ordering of it, while the re-ranker
delivers 0.2131 of that. First in seven of nine is
the honest standing on this frame, and the wall is a characterised limit that the
measurement, its method, and the levers it rules out set out in
The BP wall is a ranking limit, not an evidence ceiling.
One number, one home
This chapter states the board once. The metric definition and the
cafaeval recipe live in CAFA Evaluation Protocol; the
re-ranker design lives in ADR-D43: Stacked meta-reranker (evidence scorers plus a shallow per-category combiner); the
step-by-step reproduction path lives in
Reproduce the sealed board. Numbers are not restated elsewhere in
the book; the other chapters cross-reference this board.
Reproducing the board#
The board is produced on-platform, every result carrying the job id that
produced it, on a frame that reproduces bit-identically across two
independent runs. The ordered path (stand up the stack, load the v227
snapshot, compute the learned-encoder codes, retrieve, re-rank, and score
with cafaeval on the sealed settings) is documented in
Reproduce the sealed board, which is explicit about which stages are
job-backed today and which are not yet automated.
Provenance of the earlier figures#
An earlier version of this chapter reported an Fmax board on the GOA 220 to 229 window with ESM-C 300M embeddings and a three-generation LightGBM progression. That frame was superseded (different metric, different backbone, different window) and its numbers were never the sealed board. The full text is retained for provenance only in Pre-v227 results (superseded, retained for provenance) and must not be quoted as a current result.
See also
CAFA Evaluation Protocol: the CAFA temporal-holdout protocol, the
f_micro_wdefinition, the per-band OBO/IA registry, and the LAFA scoring parity mapping.ADR-D43: Stacked meta-reranker (evidence scorers plus a shallow per-category combiner): the stacked per-category re-ranker design.
Reproduce the sealed board: the ordered reproduction path.
Pre-v227 results (superseded, retained for provenance): the superseded pre-v227 figures, retained for provenance.