Data Scientist · Clinical Research · Bergamo, Italy
Rasoul Samei
I work on intensive-care data at Istituto Mario Negri, mostly on making fragmented clinical databases usable enough to answer a question. I'm looking for a doctoral position or a research role in health data next.
About
I'm a data scientist in clinical research. Since November 2024 I've been at the Istituto di Ricerche Farmacologiche Mario Negri in Ranica. Most of my work is intensive-care data. I harmonise fragmented clinical databases, build the analytical tables other people's studies run on, and do the statistics that go with them.
I built a dashboard with an embedded language-model assistant, so researchers here can query study data in plain language for their own work. Separately, on a residency thesis at Università di Bologna, I built a no-code extraction workbook that let the physician run the ICU analysis without a data-management background and without me in the loop.
Much of the rest is judging what the data can actually support, which I learned working next to clinicians and statisticians.
Before Mario Negri I spent six months in the UK at The Openwork Partnership, on anomaly detection for regulated financial advisory data, and before that I was a teaching assistant in machine learning at the University of Bergamo.
- 3Peer-reviewed publications
- 900+Lines of BigQuery SQL in one production pipeline
- 28+Business KPIs defined and validated
- #1Stat-Hackathon, Bergamo 2022, with my team
Research
I work on severity scores, model calibration and longitudinal analysis. But first someone has to make multi-centre clinical databases comparable enough that a question can be asked of them at all, and that is where most of my time goes. So far that has meant GiViTI and MIMIC data.
-
2025
Development and Validation of the Sequential Organ Failure Assessment (SOFA)-2 Score
JAMA, 334(23): 2090–2103
Contributing author. Extracted and processed data from the MargheritaTre (GiViTI) clinical database, statistical analysis, manuscript review.
-
2025
Anomaly Detection in Financial Advisory Services: A Machine Learning Approach for Mortgage Conduct of Business Advisers
ICSEM 2025
Unsupervised anomaly detection on regulated mortgage advisory conduct data, built in Azure Machine Learning for compliance review. Developed from my master's thesis.
-
2025
Assessing the Efficacy of a Sequential General Variational Mode Decomposition-Based Combination Model for US Wind Power Forecasting
Modeling Earth Systems and Environment
Decomposition-based hybrid time series model for wind power forecasting.
Work
-
A single validated table from a fragmented ICU database
Roughly 900 lines of SQL on Google BigQuery that turn a highly fragmented multi-table intensive-care database into one validated analytical table.
Now used as a standard data asset by other teams at the institute and by external researchers. Automating the recurring processing and reporting around it cut manual handling across several projects.
BigQuery · SQL · Python · R
-
SOFA-2 score, published in JAMA
I did the ETL on the raw MargheritaTre (GiViTI) data, which came as nested JSON and free text, calculated the SOFA-2 table, and then did the stats and analysis for the paper.
ICU mortality by total score, SOFA-1 against SOFA-2, on the first ICU day. Redrawn from Figure 3B of Ranzani et al., JAMA 2025. Contributing author. The study ran as a large international collaboration across several institutions.
R · SQL · Severity scores · Model calibration
-
A dashboard researchers can question directly
An interactive analytics dashboard with an embedded language-model assistant, so researchers can explore study data in plain language instead of asking someone to write the query for them.
Used in day-to-day research and analysis at the institute by people who do not write SQL.
Python · LangChain · SQL
-
Data engineering for an ICU nutrition study
Data engineer on a residency thesis in anaesthesiology and critical care, on nutritional therapy management in critically ill ICU patients.
I built a no-code extraction workbook so the analysis could be run directly, without a data-management background and without me in the loop.
SQL · Data extraction · ICU data
-
Anomaly detection for regulated financial advice
An unsupervised anomaly detection system in Azure Machine Learning that surfaces irregular patterns in mortgage advisory conduct data for compliance review. Alongside it, 28+ business KPIs defined, computed, validated, and delivered through Power BI dashboards to senior supervisors.
The Azure ML pipeline. New raw data is cleaned and turned into features, a warm-started model scores it, and the last step computes SHAP values and PCA plots for the Power BI dashboard the supervisors worked from. The same new data then goes back into training, so the model is ready for the next round. The model supported the evolution of the supervision function at the firm, and the method was published at ICSEM 2025. I also built Azure Python SDK and Spark SQL pipelines, and worked with a colleague on converting legacy COBOL applications into web-based services and bringing AI into selected projects.
Azure ML · Python · Spark SQL · Power BI
-
Financial Advisor, an open-source decision system
A local system that vets household money decisions against dated evidence from 14 public sources, a written investment policy, and gates that cap the score when the evidence is thin. A claim without three independent fresh sources stays undetermined instead of rounding up to a yes.
Python, standard library only, 559 tests. Open source, with the provider contract documented so someone else can add a data source or a jurisdiction. Walkthrough · Repository
Python · Public APIs · Claude Code
-
Learning projects
HL7 FHIR R4 records mapped into OMOP-style staging tables on Databricks. An NLP pipeline in spaCy feeding XGBoost classifiers for outcome prediction. A retrieval-augmented question answering assistant over a private document corpus, built with LangChain and pgvector, that keeps an audit log so every answer traces back to its source.
I built these to learn the stacks, not for a study or a client. Nobody is running them in production. They are here because they are the ground I want to work on next.
Databricks · spaCy · XGBoost · LangChain · pgvector · FastAPI · Docker · MLflow
Path
- 2024 –Data Scientist, Istituto di Ricerche Farmacologiche Mario Negri IRCCS, Ranica
- 2023 – 2024Data Scientist, six-month internship, The Openwork Partnership, Swindon, UK
- 2021 – 2024MA Economics and Data Analysis, Università degli Studi di Bergamo. Teaching assistant in Machine Learning.
- 2022First place with my team, Stat-Hackathon, Bergamo (Python, R, SAS)
- 2015BA Business Management, Islamic Azad University of Neyshabur
Toolkit
- Languages
- Python, R, SQL / Spark SQL
- Statistics
- Hypothesis testing, regression modelling, longitudinal and time series analysis, multivariate analysis
- Machine learning
- Supervised and unsupervised methods, anomaly detection, clustering, feature engineering, NLP
- Cloud & data
- Google BigQuery, Azure ML, Azure Synapse, Databricks, Power BI
- Shipping models
- MLflow, FastAPI, Docker, Git, LangChain, Claude Code
- Clinical data
- GiViTI / MargheritaTre, MIMIC, HL7 FHIR R4, OMOP
- Spoken
- Persian (native), English (professional, IELTS 7.5), Italian and German (basic)
Get in touch.
Open to research and industry roles across Europe, and to doctoral positions in health data.
rasoul@rasoulsamei.com