PROTEA#
PROtein functional Embedding-based Annotation
PROTEA is the target platform for the progressive consolidation of the Protein Information System (PIS) and FANTASIA codebases. It provides a clean, decoupled architecture for large-scale protein data ingestion, metadata enrichment, and job orchestration.
Start here Bring up the full stack from a fresh checkout and run your first job in about ten minutes.
Design System layers, job lifecycle, data model, the full operation catalogue, the CAFA evaluation protocol, and the ADRs that explain why.
autodoc Symbol-level documentation for protea.core,
protea.infrastructure, the FastAPI routers, and every worker class.
Performance Big-O profile per pipeline stage, measured hot paths, and a guide to profiling with scalene and pyinstrument.
Evidence The sealed board on the leakage-free temporal frame, with the metric definition, scoring recipe, and reproduction path cross-referenced from one home.
What is PROTEA?
A platform for protein functional annotation: from sequence ingestion through GPU embedding computation (ESM-2, ESM-C, T5/ProstT5, Ankh), a learned k-WTA retrieval encoder, KNN candidate generation, and a stacked per-category re-ranker, to board-faithful CAFA evaluation, with clean separation of infrastructure, execution flow, and domain logic.
New here? Start with the quickstart, then read the sealed board and its evidence in Results.
Documentation
- PROTEA at a glance
- Abstract
- Introduction
- Related work
- Architecture
- Computational Complexity
- Plugin authoring guide
- Results
- Reproduce the sealed board
- Appendix
- Runbooks
- Deployment Guide
- Secrets management runbook (sops + age onboarding)
- Disaster Recovery
- Stale Job Reaper
- DLQ Triage
- Ngrok Deploy Recovery
- Embedding Worker OOM
- schema_sha_v2 backfill
- schema_sha_v2 rollout (T1.6)
- Observability: OpenTelemetry SDK
- Observability Operator Runbook
- Observability: Loki log aggregation
- Observability: Prometheus metrics
- Process-Based Stack Deployment Guide
- LAFA Native Parity (INT-8 prep)
- Serve Learned-Code Retrieval (novel queries)
- Clean Reproducible Evaluation Frame (R0.1)
- Release process: trunk + snapshot promotion
- Resumable export + parallel minijobs
- Quality Engineering
- Operational Insights and Lessons Learned
- Glossary
- References