Abstract
Machine learning (ML) offers promising opportunities to accelerate hydropower scheduling, yet its effectiveness is constrained by the limited size and variability of historical training data. This paper presents a scalable methodology for generating synthetic datasets that remain physically consistent and economically representative. The approach constructs a multivariate kernel density estimation (KDE) model to sample correlated initial reservoir states, which are then linked to historical data to preserve seasonal characteristics. Inflow and electricity price scenarios are subsequently sampled from historical archives to ensure realistic hydrological patterns and market conditions. Water values are derived using a normalized availability coefficient adjusted for both synthetic reservoir states and sampled price profiles. Applied to the Tokke-Vinje watercourse in Norway, the method expands approximately 3,650 historical cases to an arbitrary number of synthetic scenarios. The resulting dataset enables the development of robust ML models for hydropower scheduling and supports their integration into the optimization tool.