Berlin, Germany Max Delbrück Center

Katarína Grešová

Portrait of Katarína Grešová

Machine learning for biology, built to be reused

My work is the loop from experiment to model to biological insight: benchmarks and data preparation at one end, model interpretation at the other, models of RNA regulation in between.

Training a good model is hard, and it is the part that gets all the attention. But on flawed data, or with no way to see what it learned, it still isn’t usable — and if it isn’t usable, why did we build it?

I was a software engineer before I was a scientist — four companies, Oracle and Tieto among them — and I have learned a new field every few years since. What I want next is a company where that whole combination is the job: owning a research question, being the one who checks whether the answer is real, and explaining both to the people who weren’t in the room.

Looking for a research scientist or project lead role. Bio, health and techbio — Berlin or remote in Europe.

Start a conversation

I am a computational biologist who was a software engineer first. That combination shows up in three ways, and every one of them is something I have already done for somebody.

  • Own the question

    A biological question taken from raw data to a model to an answer somebody can act on.

    Five fields so far — miRNA targeting, RNA-binding proteins, protein structure, senescence, translation — four of them entered knowing nothing. Work out what the data can honestly answer, build the smallest thing that answers it, and hand over the code that regenerates the result.

    Recent work in Molecular Cell and Nucleic Acids Research. Doctorate awarded summa cum laude. See it

  • Find out whether it is real

    The person who notices the model is right for the wrong reason — before a partner or a regulator does.

    The dataset a classifier can pass without learning anything, leakage between train and test, the number that is real but does not mean what the sentence around it says. That is the difference between a result you can put in front of a partner and one that falls apart on their data — and I would rather build the check into CI than do it by hand twice.

    Packaged as Genomic Benchmarks QC — talks at EMBO and RECOMB-Seq this year. The benchmark suite behind it is cited 175 times. See it

  • Sit between the science and everyone else

    I explain the biology to the engineers and the modelling to the biologists, and keep the work moving while I do it.

    Deep learning taught to biologists who had never written a training loop; molecular biology to computer scientists. Students and interns supervised, hackathons run, a project led across two institutions where students built most of it. This is the part I want more of, not less.

    Two of the courses are public and still runnable, so you can watch me explain something before we ever talk. See it

11
courses and workshops taught, students supervised
316
citations across 16 papers, 5 of them mine to lead
182
stars on the benchmark suite, cited 175 times
4
companies whose production software I helped ship

Citation and star counts checked September 2026.

Everything I've built

Everything here was used by somebody who isn't me — that is the only rule for getting on the list. Of the 14, 3 began as a complaint about an evaluation and 5 were shipped in industry, where somebody else's day was ruined if they broke.

All publications

  1. 2026

    SenCat: cataloging human cell senescence through multi-omic profiling of multiple senescent primary cell types

    Carlos Anerillas, Gisela Altés, Katarína Grešová, Dimitrios Tsitsipatis, Krystyna Mazan-Mamczarz, et al., Manolis Maragkakis, Nathan Basisty, Myriam Gorospe

    Molecular Cell86(13), 2605–2616.e8Shared second author

  2. 2025

    miRBench: novel benchmark datasets for microRNA binding site prediction that mitigate against prevalent microRNA frequency class bias

    Stephanie Sammut, Katarína Grešová, Dimosthenis Tzimotoudis, Eva Maršálková, David Čechák, Panagiotis Alexiou

    Bioinformatics41(Supplement_1), i542–i551Co-first author

  3. 2024

    RBP-Tar: a searchable database for experimental RBP binding sites

    Katarína Grešová, Tomáš Racek, Vlastimil Martinek, David Čechák, Radka Svobodová, Panagiotis Alexiou

    F1000Research12, 755First author

  4. 2023

    Genomic benchmarks: a collection of datasets for genomic sequence classification

    Katarína Grešová, Vlastimil Martinek, David Čechák, Petr Šimeček, Panagiotis Alexiou

    BMC Genomic Data24, 25First author

  5. 2023

    Using attribution sequence alignment to interpret deep learning models for miRNA binding site prediction

    Katarína Grešová, Ondřej Vaculík, Panagiotis Alexiou

    Biology12(3), 369First author

Not hiring, but have a group that needs this done rather than staffed? These three are bookable on their own, and I have done each of them before.

  • Teach your group

    The computational skills a science degree leaves out, taught to the people who need them on Monday.

    The Missing Skills is the written version, open and free. The three-day deep learning course at the University of Malta took a room of biologists from never having written a training loop to having one that worked.

    See it
  • Set it up with you

    The part where good practice survives contact with your actual project.

    Two published papers rest on workflows I wrote, so their figures can be regenerated from raw data by someone who was not in the room. Before the doctorate I spent two years on nothing but continuous integration and test infrastructure.

    See it
  • Take on the problem

    Bring me the question nobody has had time to get to. I work out whether your data can answer it, and build the first version that shows whether it is worth pursuing.

    Most of my career has been exactly this: arriving in somebody else's group, taking on a question that was already theirs, and leaving a published answer behind. Five groups, four countries — the results are in Molecular Cell, Nucleic Acids Research and Bioinformatics.

    See it

Let's talk

If you are building something in biology or health and you need someone to own a question, keep the results honest, and explain both to the rest of the company — that is the conversation I am quickest to answer. You don't need a worked-out role description; a paragraph about what you're building is plenty. Teaching and consulting enquiries are welcome too.