Data Scientist · Clinical Research · Bergamo, Italy

Rasoul Samei

I work on intensive-care data at Istituto Mario Negri, mostly on making fragmented clinical databases usable enough to answer a question. I'm looking for a doctoral position or a research role in health data next.

Rasoul Samei

I'm a data scientist in clinical research. Since November 2024 I've been at the Istituto di Ricerche Farmacologiche Mario Negri in Ranica. Most of my work is intensive-care data. I harmonise fragmented clinical databases, build the analytical tables other people's studies run on, and do the statistics that go with them.

I built a dashboard with an embedded language-model assistant, so researchers here can query study data in plain language for their own work. Separately, on a residency thesis at Università di Bologna, I built a no-code extraction workbook that let the physician run the ICU analysis without a data-management background and without me in the loop.

Much of the rest is judging what the data can actually support, which I learned working next to clinicians and statisticians.

Before Mario Negri I spent six months in the UK at The Openwork Partnership, on anomaly detection for regulated financial advisory data, and before that I was a teaching assistant in machine learning at the University of Bergamo.

  • 3Peer-reviewed publications
  • 900+Lines of BigQuery SQL in one production pipeline
  • 28+Business KPIs defined and validated
  • #1Stat-Hackathon, Bergamo 2022, with my team

I work on severity scores, model calibration and longitudinal analysis. But first someone has to make multi-centre clinical databases comparable enough that a question can be asked of them at all, and that is where most of my time goes. So far that has meant GiViTI and MIMIC data.

  1. 2025
    Development and Validation of the Sequential Organ Failure Assessment (SOFA)-2 Score

    Ranzani OT, Singer M, et al., for the SOFA-2 Collaborators, including Rasoul Samei

    JAMA, 334(23): 2090–2103

    Contributing author. Extracted and processed data from the MargheritaTre (GiViTI) clinical database, statistical analysis, manuscript review.

  2. 2025
    Anomaly Detection in Financial Advisory Services: A Machine Learning Approach for Mortgage Conduct of Business Advisers

    Rasoul Samei

    ICSEM 2025

    Unsupervised anomaly detection on regulated mortgage advisory conduct data, built in Azure Machine Learning for compliance review. Developed from my master's thesis.

  3. 2025
    Assessing the Efficacy of a Sequential General Variational Mode Decomposition-Based Combination Model for US Wind Power Forecasting

    Shalchilar M, Samei R, et al.

    Modeling Earth Systems and Environment

    Decomposition-based hybrid time series model for wind power forecasting.

  • A single validated table from a fragmented ICU database

    Mario Negri · 2024–present

    Roughly 900 lines of SQL on Google BigQuery that turn a highly fragmented multi-table intensive-care database into one validated analytical table.

    Now used as a standard data asset by other teams at the institute and by external researchers. Automating the recurring processing and reporting around it cut manual handling across several projects.

    BigQuery · SQL · Python · R

  • SOFA-2 score, published in JAMA

    Mario Negri · 2025

    I did the ETL on the raw MargheritaTre (GiViTI) data, which came as nested JSON and free text, calculated the SOFA-2 table, and then did the stats and analysis for the paper.

    0 20 40 60 80 100 0 4 8 12 16 20 23/24 Total SOFA points on ICU day 1 ICU mortality, % SOFA-2 SOFA-1
    ICU mortality by total score, SOFA-1 against SOFA-2, on the first ICU day. Redrawn from Figure 3B of Ranzani et al., JAMA 2025.

    Contributing author. The study ran as a large international collaboration across several institutions.

    R · SQL · Severity scores · Model calibration

  • A dashboard researchers can question directly

    Mario Negri · 2025

    An interactive analytics dashboard with an embedded language-model assistant, so researchers can explore study data in plain language instead of asking someone to write the query for them.

    Used in day-to-day research and analysis at the institute by people who do not write SQL.

    Python · LangChain · SQL

  • Data engineering for an ICU nutrition study

    With Università di Bologna

    Data engineer on a residency thesis in anaesthesiology and critical care, on nutritional therapy management in critically ill ICU patients.

    I built a no-code extraction workbook so the analysis could be run directly, without a data-management background and without me in the loop.

    SQL · Data extraction · ICU data

  • Anomaly detection for regulated financial advice

    The Openwork Partnership, Swindon UK · 2023–2024

    An unsupervised anomaly detection system in Azure Machine Learning that surfaces irregular patterns in mortgage advisory conduct data for compliance review. Alongside it, 28+ business KPIs defined, computed, validated, and delivered through Power BI dashboards to senior supervisors.

    Cleaning new raw data Feature engineering derived features Scoring trained model on new data Explaining SHAP values PCA plots Training warm start on new data raw conduct data training dataset Power BI dashboard
    The Azure ML pipeline. New raw data is cleaned and turned into features, a warm-started model scores it, and the last step computes SHAP values and PCA plots for the Power BI dashboard the supervisors worked from. The same new data then goes back into training, so the model is ready for the next round.

    The model supported the evolution of the supervision function at the firm, and the method was published at ICSEM 2025. I also built Azure Python SDK and Spark SQL pipelines, and worked with a colleague on converting legacy COBOL applications into web-based services and bringing AI into selected projects.

    Azure ML · Python · Spark SQL · Power BI

  • Financial Advisor, an open-source decision system

    Personal project · MIT

    A local system that vets household money decisions against dated evidence from 14 public sources, a written investment policy, and gates that cap the score when the evidence is thin. A claim without three independent fresh sources stays undetermined instead of rounding up to a yes.

    Python, standard library only, 559 tests. Open source, with the provider contract documented so someone else can add a data source or a jurisdiction. Walkthrough · Repository

    Python · Public APIs · Claude Code

  • Learning projects

    Ongoing

    HL7 FHIR R4 records mapped into OMOP-style staging tables on Databricks. An NLP pipeline in spaCy feeding XGBoost classifiers for outcome prediction. A retrieval-augmented question answering assistant over a private document corpus, built with LangChain and pgvector, that keeps an audit log so every answer traces back to its source.

    I built these to learn the stacks, not for a study or a client. Nobody is running them in production. They are here because they are the ground I want to work on next.

    Databricks · spaCy · XGBoost · LangChain · pgvector · FastAPI · Docker · MLflow

Languages
Python, R, SQL / Spark SQL
Statistics
Hypothesis testing, regression modelling, longitudinal and time series analysis, multivariate analysis
Machine learning
Supervised and unsupervised methods, anomaly detection, clustering, feature engineering, NLP
Cloud & data
Google BigQuery, Azure ML, Azure Synapse, Databricks, Power BI
Shipping models
MLflow, FastAPI, Docker, Git, LangChain, Claude Code
Clinical data
GiViTI / MargheritaTre, MIMIC, HL7 FHIR R4, OMOP
Spoken
Persian (native), English (professional, IELTS 7.5), Italian and German (basic)

Get in touch.

Open to research and industry roles across Europe, and to doctoral positions in health data.

rasoul@rasoulsamei.com