Sagace

Sagace

Models for living systems

A researcher looking across a lake toward a mountain.
About

AI systems for biology and life.

Sagace is a research lab exploring how models can represent, understand, and reason about living systems.

Can models reason through living systems?

Biology is difficult to understand and experiments are expensive. Palaestra tests whether models can interpret evidence, form hypotheses, work through failure, and choose the next useful experiment.

pass@1pass@4pass@8
Illustration of biological reinforcement learning environments
Benchmark suite — for frontier labs

Palaestra

Verifiable environments for biological reasoning.

A curated suite of verifiable tasks that train and evaluate models on real biological, chemical, and pharmaceutical reasoning. Each task is usable interoperably — as a reinforcement-learning environment or as a benchmark — with a hidden oracle and a runnable verifier.

24 puzzle boxes192+ verifiable tasksContact us ↗
Puzzle boxes

Themed collections of related tasks — one box is one benchmark. Public boxes may be used for training; private boxes are held out for evaluation.

Cells & single-cell

  • 01Cell reasoning & identityAnnotate cell type, origin and disease state under distribution shift, with calibrated abstention.scRNA-seq
  • 02Batch-aware processingIntegrate away technical batch effects while conserving biology, judged on a held-out batch.scRNA-seq
  • 03QC & outlier handlingFlag hidden corrupted cells without discarding rare, genuine populations.scRNA-seq
  • 04Missing-modality imputationImpute a masked modality — RNA, protein or ATAC — with calibrated per-entry uncertainty.multiome
  • 05AnonymizationDefeat a hidden re-identification attacker while preserving downstream analysis on the privacy–utility frontier.scRNA-seq

Omics & mechanism

  • 06Multi-omics integrationFuse RNA, protein and ATAC to predict a hidden held-out phenotype.multi-omics
  • 07Causal perturbation inferenceRank causal genes from perturbation data and propose the next intervention.Perturb-seq
  • 08Pathway & mechanism inferenceRecover the truly perturbed pathways and regulators, with false-discovery control.transcriptomics
  • 09Variant interpretationRank variants by functional or clinical effect on gene-disjoint splits that block memorization.DNA · protein

Proteins & structure

  • 10Protein variant designDesign variants under mutation and developability budgets, scored for fitness, stability and diversity.sequence
  • 11Structural bioinformaticsIdentify binding sites and rank plausible poses against hidden structural ground truth.structure

Small molecules & chemistry

  • 12Small-molecule designReferenceMulti-parameter optimization of drug-like molecules, graded by RDKit validity plus hidden property oracles.SMILES
  • 13Lead prioritizationRank candidate series on potency, selectivity, ADMET, novelty and synthesizability.SMILES
  • 14Assay triage & active learningChoose which compounds to test next as labels are revealed round by round.SMILES

Pharmacology & clinic

  • 15Dose–response & PK/PDRecover EC50, Hill slope and clearance and predict held-out dose/time points.timeseries
  • 16Drug portfolio managementAdvance, pause or kill programs under a budget against a stochastic outcome simulator.simulator
  • 17Cohort constructionBuild a temporally valid, leakage-free cohort from hidden patient records.EHR · OMOP
  • 18Epidemiological inferenceEstimate R0 and intervention effects from synthetic outbreak curves.timeseries

Method, data & rigor

  • 19Experiment & assay designPropose designs a hidden simulator scores for statistical power, confound control and cost.simulator
  • 20Metadata reconciliationHarmonize heterogeneous samples to canonical ontology labels under consistency constraints.ontology
  • 21Data-leakage auditDetect and patch hidden train/test overlap and target leakage in an ML pipeline.code
  • 22ReproducibilityRe-run a batch of computational experiments and reproduce their claimed results.code
  • 23Literature claim verificationLabel a claim supported, refuted or insufficient, citing the exact evidence passages.text
  • 24Dual-use & safety triageClassify whether a request crosses a hidden safety threshold — defensive evaluation only.text
The work ahead

Progress in biology should be something you can prove.

Models will reason about living systems only as well as we can measure them. Palaestra builds the verifiable environments to train, evaluate, and trust that reasoning — one task at a time, with a hidden oracle behind every one.

Training frontier models for biology, or evaluating them? Palaestra is opening to research partners.

The landscape

Benchmarks in biology

A field guide to the public benchmarks the community already uses to measure models on biological reasoning — the terrain Palaestra is built to extend.

  • 01Proteins

    ProteinGym

    Mutation-effect and fitness prediction across 250+ deep mutational scanning assays and curated clinical variants.

    Notin et al. · 2023
  • 02Proteins

    TAPE

    Tasks Assessing Protein Embeddings

    Five biologically grounded tasks measuring how well learned protein representations transfer.

    Rao et al. · 2019
  • 03Proteins

    CASP

    Critical Assessment of Structure Prediction

    The biennial blind experiment assessing prediction of protein 3D structure from sequence.

    Prediction Center · 1994–
Products for biologists

Predictions you can interrogate.

Palaestra is built for the labs training the models. These are built for the scientists at the bench: each takes a compound and returns more than a number — the pathways that moved, the evidence behind them, and an honest account of where the model stops knowing. All three run in your browser.

Product access — for scientistsPoint these models at your own molecules

Both demos run the shipped models against a fixed compound library. Tell us what you are screening and we will open them up to your series — and to the full technical reports behind each call.