Berlin, Germany Max Delbrück Center
Katarína Grešová
Machine learning for biology, built to be reused
My work is the loop from experiment to model to biological insight: benchmarks and data preparation at one end, model interpretation at the other, models of RNA regulation in between.
Training a good model is hard, and it is the part that gets all the attention. But on flawed data, or with no way to see what it learned, it still isn’t usable — and if it isn’t usable, why did we build it?
I was a software engineer before I was a scientist — four companies, Oracle and Tieto among them — and I have learned a new field every few years since. What I want next is a company where that whole combination is the job: owning a research question, being the one who checks whether the answer is real, and explaining both to the people who weren’t in the room.
Looking for a research scientist or project lead role. Bio, health and techbio — Berlin or remote in Europe.
What I bring to a team Start a conversation
I am a computational biologist who was a software engineer first. That combination shows up in three ways, and every one of them is something I have already done for somebody.
-
Own the question
A biological question taken from raw data to a model to an answer somebody can act on.
Five fields so far — miRNA targeting, RNA-binding proteins, protein structure, senescence, translation — four of them entered knowing nothing. Work out what the data can honestly answer, build the smallest thing that answers it, and hand over the code that regenerates the result.
Recent work in Molecular Cell and Nucleic Acids Research. Doctorate awarded summa cum laude. See it
-
Find out whether it is real
The person who notices the model is right for the wrong reason — before a partner or a regulator does.
The dataset a classifier can pass without learning anything, leakage between train and test, the number that is real but does not mean what the sentence around it says. That is the difference between a result you can put in front of a partner and one that falls apart on their data — and I would rather build the check into CI than do it by hand twice.
Packaged as Genomic Benchmarks QC — talks at EMBO and RECOMB-Seq this year. The benchmark suite behind it is cited 175 times. See it
-
Sit between the science and everyone else
I explain the biology to the engineers and the modelling to the biologists, and keep the work moving while I do it.
Deep learning taught to biologists who had never written a training loop; molecular biology to computer scientists. Students and interns supervised, hackathons run, a project led across two institutions where students built most of it. This is the part I want more of, not less.
Two of the courses are public and still runnable, so you can watch me explain something before we ever talk. See it
- 11
- courses and workshops taught, students supervised
- 316
- citations across 16 papers, 5 of them mine to lead
- 182
- stars on the benchmark suite, cited 175 times
- 4
- companies whose production software I helped ship
Citation and star counts checked September 2026.
Things I build Everything I've built
Everything here was used by somebody who isn't me — that is the only rule for getting on the list. Of the 14, 3 began as a complaint about an evaluation and 5 were shipped in industry, where somebody else's day was ruined if they broke.
-
Genomic Benchmarks QC
Automated quality control for genomic machine-learning datasets. It scores the biases, duplicate sequences and train/test leakage a classifier could exploit before you train on it — length differences, GC content, per-position give-aways, near-duplicate overlap...
-
miRBench
Benchmark datasets and a Python package for miRNA target-site prediction, built to remove the frequency-class bias that quietly inflates published scores. It ships the data, the splits and wrappers around the existing predictors, so...
-
Genomic Benchmarks
Eight curated datasets for genomic sequence classification across human, mouse and roundworm, each with a sensible split and a baseline model, so a new architecture has something honest to beat. Installable as a package...
-
Attribution sequence alignment
A way to make a model say what it noticed. Attribution gives you one importance score per nucleotide, which is not yet a motif; aligning those profiles across many sequences makes the patterns a...
-
AlphaFind and AlphaFind2
Structure-similarity search across the whole of AlphaFold DB. Pretrained networks compress each structure into an embedding, which turns an all-against-all comparison over more than 200 million structures into a query you can wait for....
-
RBP-Tar
A searchable database of experimentally determined RNA-binding protein sites, pulled out of published CLIP experiments and put behind a query interface — so you can ask what binds a transcript without reprocessing somebody else's...
Teaching Everything I've taught
I teach in both directions — computation to biologists, biology to computer scientists — in university practicals, conference tutorials, a three-day deep learning course, and material anyone can work through on their own. Most of it is public, which makes it the easiest way to check whether I can explain something before you hire me to.
-
Deep Learning for Genomics
Convolutional and recurrent models for biological sequence data, taught from first principles to a room of biologists who had not written a training loop before, and who had one working by the end of...
-
The Missing Skills
Project-based tutorials on the computational skills a science degree leaves out — version control, the shell, and making your work runnable by someone who isn't you. Written because the bottleneck in computational biology is...
-
Biology Crash Course
Molecular biology from the ground up, written for computer scientists who need enough of it to be dangerous. The mirror image of The Missing Skills, and between them the two directions I spend most...
Selected publications All publications
-
2026
SenCat: cataloging human cell senescence through multi-omic profiling of multiple senescent primary cell types
Molecular Cell86(13), 2605–2616.e8Shared second author
-
2025
miRBench: novel benchmark datasets for microRNA binding site prediction that mitigate against prevalent microRNA frequency class bias
Bioinformatics41(Supplement_1), i542–i551Co-first author
-
2024
RBP-Tar: a searchable database for experimental RBP binding sites
F1000Research12, 755First author
-
2023
Genomic benchmarks: a collection of datasets for genomic sequence classification
BMC Genomic Data24, 25First author
-
2023
Using attribution sequence alignment to interpret deep learning models for miRNA binding site prediction
Biology12(3), 369First author
Also available for
Not hiring, but have a group that needs this done rather than staffed? These three are bookable on their own, and I have done each of them before.
-
Teach your group
The computational skills a science degree leaves out, taught to the people who need them on Monday.
The Missing Skills is the written version, open and free. The three-day deep learning course at the University of Malta took a room of biologists from never having written a training loop to having one that worked.
See it -
Set it up with you
The part where good practice survives contact with your actual project.
Two published papers rest on workflows I wrote, so their figures can be regenerated from raw data by someone who was not in the room. Before the doctorate I spent two years on nothing but continuous integration and test infrastructure.
See it -
Take on the problem
Bring me the question nobody has had time to get to. I work out whether your data can answer it, and build the first version that shows whether it is worth pursuing.
Most of my career has been exactly this: arriving in somebody else's group, taking on a question that was already theirs, and leaving a published answer behind. Five groups, four countries — the results are in Molecular Cell, Nucleic Acids Research and Bioinformatics.
See it
Let's talk
If you are building something in biology or health and you need someone to own a question, keep the results honest, and explain both to the rest of the company — that is the conversation I am quickest to answer. You don't need a worked-out role description; a paragraph about what you're building is plenty. Teaching and consulting enquiries are welcome too.