PROTEA#

PROtein functional Embedding-based Annotation

PROTEA is the target platform for the progressive consolidation of the Protein Information System (PIS) and FANTASIA codebases. It provides a clean, decoupled architecture for large-scale protein data ingestion, metadata enrichment, and job orchestration.

Quickstart

Start here Bring up the full stack from a fresh checkout and run your first job in about ten minutes.

Installation and Quickstart
Architecture

Design System layers, job lifecycle, data model, the full operation catalogue, the CAFA evaluation protocol, and the ADRs that explain why.

Architecture
API Reference

autodoc Symbol-level documentation for protea.core, protea.infrastructure, the FastAPI routers, and every worker class.

API Reference
Complexity

Performance Big-O profile per pipeline stage, measured hot paths, and a guide to profiling with scalene and pyinstrument.

Computational Complexity
Results

Evidence The sealed board on the leakage-free temporal frame, with the metric definition, scoring recipe, and reproduction path cross-referenced from one home.

Results

What is PROTEA?

A platform for protein functional annotation: from sequence ingestion through GPU embedding computation (ESM-2, ESM-C, T5/ProstT5, Ankh), a learned k-WTA retrieval encoder, KNN candidate generation, and a stacked per-category re-ranker, to board-faithful CAFA evaluation, with clean separation of infrastructure, execution flow, and domain logic.

New here? Start with the quickstart, then read the sealed board and its evidence in Results.