Dataset
Dataset of Korea Clinical Datathon 2026
K-MIMIC
K-MIMIC (Korea Medical Information Mart for Intensive Care) is an extensive health-related database derived from patients who were admitted to the intensive care units (ICUs) of hospitals in Korea. This database has been established through the support of a grant from the Korea Health Technology R&D Project, managed by the Korea Health Industry Development Institute (KHIDI), and funded by the Ministry of Health & Welfare, Republic of Korea (grant number: HI21C1074). K-MIMIC comprises a wide range of data, including electronic medical records, medical images, and vital sign information, serving as a crucial resource for research and development in the field of intensive care medicine.
Reference link: https://sites.google.com/view/k-mimic
MIMIC-IV
The MIMIC-IV dataset is an extensive health-related database collected from patients admitted to the intensive care units (ICUs) at the Beth Israel Deaconess Medical Center (BIDMC) between 2008 and 2019. It includes a wide range of data such as patient demographics, vital signs, laboratory test results, procedures, medications, and outcomes. This dataset is instrumental for research in critical care medicine, offering detailed insights into patient care and clinical outcomes in the ICU.
Reference link: https://mimic.mit.edu/docs/
VitalDB
The dataset was obtained from non-cardiac surgery patients, including those undergoing general, thoracic, urologic, and gynecologic procedures, who had either routine or emergency surgeries. These surgeries were conducted in 10 operating rooms over the course of one year, from August 2016 to July 2017, at a single academic hospital, Seoul National University Hospital, located in Seoul, Republic of Korea. This dataset provides a comprehensive overview of surgical practices and patient outcomes in a high-volume, academic medical center.
Reference link: https://vitaldb.net/dataset/
INSPIRE
The INSPIRE dataset is a comprehensive collection of perioperative medical data from adult surgical cases at Seoul National University Hospital (SNUH) between 2011 and 2020. This dataset encompasses 131,109 cases involving 99,900 patients aged between 18 and 90, all of whom underwent various surgical procedures under anesthesia. It provides extensive information that includes patient demographics, surgical details, and anesthesia records, serving as a valuable resource for research in perioperative medicine and anesthesiology.
Reference link: https://physionet.org/content/inspire/1.2/
SNUH CDM
While Electronic Medical Records (EMR) and various healthcare information systems are being integrated to optimize patient care and improve treatment outcomes, effectively analyzing and interpreting complex medical data remains a challenging task. The Seoul National University Hospital (SNUH) Common Data Model (CDM) is a standardized and anonymized dataset built from EMR data to facilitate effective medical data analysis and the application of advanced machine learning technologies. This CDM is based on the OMOP (Observational Medical Outcomes Partnership) CDM standard, ensuring interoperability with data from other medical institutions and enhancing multi-center research. The dataset includes comprehensive clinical data from over 3.69 million patients, including diagnostic information, prescription data, laboratory results, and procedural information, between October 2004 and December 2023.
Reference link: https://khdp.net/database/data-search-detail/665/SNUH-CDM/1.0.0
AMC-MFC Datathon Subset
AMC-MFC Datathon Subset is drawn from AMC-MFC (AMC Multimodal Frozen Cohort), a death-confirmed cohort of approximately 245,673 patients whose death was verified at Asan Medical Center from the hospital's opening through December 31, 2025, which provides a complete clinical course to a definite endpoint without right-censoring. The subset comprises approximately 1,000 of these patients who also have operating-room vital sign recordings, drawn by stratified random sampling on sex, age group, and primary diagnosis category. It pairs structured clinical data in the OMOP Common Data Model — demographics, observation periods, visits, conditions, drug exposures, procedures, selected measurements including vital signs, and death records, together with the OHDSI standard vocabularies — with raw intraoperative waveforms (electrocardiography, arterial blood pressure, photoplethysmography, respiration) and their surgery start/end times, device names, and channel names. Direct identifiers are removed, patient and provider identifiers are pseudonymized, all dates are shifted by a per-patient random offset that preserves within-patient intervals and time-to-death, and ages of 90 years and above are top-coded; the subset is accessible only within the KHDP Credentialed Secure environment for the duration of the datathon (October 16–18, 2026) and is destroyed thereafter.
CRC-LoT
CRC-LoT (Colorectal Cancer Line of Therapy) is a benchmark dataset for reconstructing lines of therapy (LoTs) from systemic anticancer therapy (SACT) prescription records. It consists of prescription records from 3,000 patients with colorectal cancer who received anticancer therapy at Asan Medical Center between June 1997 and October 2025. Each row-level record contains the order and end dates, drug code, product name, ingredient name, cycle day, and prescribed duration. The prediction target is the complete treatment pathway for each patient: an ordered sequence of LoTs, each defined by its start and end dates, regimen, and line number. Because LoTs are not explicitly recorded in prescription records, the reference labels were established through consensus between clinical pharmacists and medical oncologists. Labels are available for the development set, while test-set labels are withheld for evaluation.