Forward-Deployed Cheminformatician
Apheris
Seniority
Midweight
Model
In-Office
Sector
Salary
Undisclosed
Contract
Full-Time
Own how binding data is prepared across co-folding focused networks and initiatives. Binding data arrives from pharma partners in heterogeneous shapes — different assay registries, metadata, chemical-representation standards, and qualifiers. You will define a repeatable, well-documented preparation pipeline that pharma representatives can run alongside Apheris, and scale it to the public-data corpus for model training. This is half engineering, half forward-deployed work.
What you'll do
- Define and own the binding-data preparation protocol — data schema, small-molecule standardization, assay metadata model, value handling (KD, Ki, IC50, pIC50), qualifier and censored-value handling, duplicate and replicate aggregation.
- Build the tooling that runs it — modular scripts, validators with actionable errors, and reusable pipelines that survive different pharma upstream systems (Dotmatics, Spotfire, in-house registries).
- Work forward-deployed with pharma. Sit with their biologists and medicinal chemists, walk them through the protocol, sense-check what an assay column actually measures, and unblock retrieval.
- Maintain the small-molecule representation pipeline — RDKit standardization, tautomer and ionization handling, stereochemistry preservation, and PAINS / frequent-hitter filtering.
- Curate the public binding-data foundation — ChEMBL, BindingDB, PubChemBioAssay — prepared to the same standard for model training.
- Hand the productized pipeline cleanly to engineering for scaling, and partner with ML to keep the data contract valid as models and networks evolve.
What you'll need
- BSc, MSc, PhD or equivalent in cheminformatics, computational chemistry, or a related field, plus 3+ years preparing biological assay data in a discovery setting.
- Fluent in Python and RDKit. SMILES normalization, tautomer / ionization / stereochemistry handling, and scaffold extraction are second nature.
- Hands-on experience curating quantitative binding assay data (KD, Ki, IC50, pIC50) and HTS data — censored values, qualifiers, duplicates, replicate aggregation, and assay metadata interpretation.
- Write good engineering code — version control, tested modular scripts, validators that return useful errors.
- Comfortable forward-deployed with pharma medicinal chemists and biologists. Can interpret column labels and encode that back into the protocol.
- Enjoy turning a messy ad-hoc cleaning job into a repeatable protocol others can run.
Nice to have
- Practical familiarity with public binding-data sources (ChEMBL, BindingDB, PubChemBioAssay) and the gotchas in each.
- Applied LLM tooling (Claude, Codex, Cursor) to accelerate data cleaning or metadata harmonization.
- Worked across institutional data boundaries — federated, multi-party, or otherwise — where the data-preparation contract has to hold under partial visibility.
- Publication record or open-source contributions in cheminformatics or quantitative pharmacology.
What they offer
- Industry-competitive compensation, including early-stage virtual share options
- Remote-first work
- Wellbeing budget, mental health support, work-from-home budget, co-working stipend, and learning budget
- Generous holiday allowance
- Office Days at Berlin HQ or different European location (3x per year)

