Reproduce the sealed board#
This page names the ordered path to reproduce PROTEA’s sealed board (the
f_micro_w headline stated in Results). It is deliberately honest
about which stages are automated on-platform today and which are not. It does
not restate the number and it does not print commands that have not been
verified against the codebase; each stage links to the runbook or ADR that
carries the exact, verified payloads.
There is no single-command reproduction
The board is produced stage by stage, every result carrying the job id that produced it, on a frame that reproduces bit-identically across two independent runs (see Clean Reproducible Evaluation Frame (R0.1)). There is no one orchestrator job that runs all stages end-to-end, and the re-ranker itself is trained out of this repository (see step 5). Treat the steps below as the map, not as a script.
The ordered path#
Stand up the stack. Bring up the API and workers from a checkout with
bash scripts/manage.sh start(Postgres and RabbitMQ must already be running). Installation and first-run details are in Installation and Quickstart. All work below is dispatched as jobs viaPOST /jobswith an{operation, payload}body; never use ad-hoc curl against internal endpoints.Load the v227 snapshot. Load the GO ontology snapshot and the GOA annotation sets for the frame with the
load_ontology_snapshotandload_goa_annotationsoperations. The canonical OBO for band v227 isreleases/2025-07-22and the IA artefact is the t0 IA for that band; the authoritative per-band OBO and IA pins are in the band registry, documented in CAFA Evaluation Protocol (per-band registry) and enforced at runtime and in CI.Compute the learned-encoder codes. The champion is the k-WTA retrieval encoder (config
d8979601), which stores GO-aligned codes rather than a raw PLM vector. Codes are materialised over the base embeddings by theapply_learned_encoderoperation offline, and a novel query is embedded on the fly bycompute_embeddingswhen its config uses thelearned-codebackend. Serving requires the head artifact to be provided through thePROTEA_LEARNED_ENCODER_ARTIFACT(orPROTEA_LEARNED_ENCODER_DIR) environment variable; the exact resolution rules and failure modes are in Serve Learned-Code Retrieval (novel queries).Retrieve. Run KNN GO transfer over the learned codes with the
predict_go_termsoperation, producing aPredictionSet. The retrieval and prediction operations and their payload schemas are in Operations.Re-rank. The candidates are re-ranked by a stacked per-category re-ranker (evidence scorers plus a shallow per-category combiner), designed in ADR-D43: Stacked meta-reranker (evidence scorers plus a shallow per-category combiner). Booster training is not part of PROTEA: it runs in the
protea-reranker-labsibling repository over a frozen parquet dataset that PROTEA publishes viaexport_research_dataset. The trained booster is brought back into PROTEA through thePOST /reranker-models/importendpoint (or thescripts/register_reranker.pyhelper). This step is therefore only partly automated on-platform: the dataset export and the import are jobs and endpoints here, the training is not.Score with cafaeval on the sealed settings. Generate the cross-OBO evaluation set with
generate_evaluation_setand score the re-rankedPredictionSetwithrun_cafa_evaluation, using the board-faithful recipe: bandv227,prop=fill,norm=cafa,no_orphans,th_step=0.01,max_terms=None, the t0 IA file, the release terms-of-interest file, and-knownon PK only. The exact job payloads, the asymmetric cross-OBO pins, and the bit-identical reproduction check are in Clean Reproducible Evaluation Frame (R0.1); the metric definition and the LAFA parity mapping are in CAFA Evaluation Protocol.
What is not yet automated#
There is no single end-to-end job that runs all six stages; each is dispatched and its output UUID threaded into the next.
Re-ranker training lives in
protea-reranker-lab(step 5), outside this repository. PROTEA automates the dataset export and the booster import, not the training.The IA and terms-of-interest artefacts (step 6) and the learned-encoder head artifact (step 3) are supplied by path or environment variable; they are not fetched automatically.
See also
Results: the sealed board this page reproduces.
Clean Reproducible Evaluation Frame (R0.1): the verified job payloads and the bit-identical reproduction check.
Serve Learned-Code Retrieval (novel queries): the learned encoder
d8979601and its head-artifact configuration.CAFA Evaluation Protocol: the metric definition, the per-band OBO/IA registry, and the cafaeval recipe.
ADR-D43: Stacked meta-reranker (evidence scorers plus a shallow per-category combiner): the re-ranker design.
Reproduction guide (superseded, retained for provenance): the superseded pre-v227 guide, retained for provenance only.