๐ Datasets โ Awesome Time Series
Loading datasetsโฆ
Loading datasetsโฆ
1,321 datasets & benchmarks โ 15 canonical foundations plus emerging datasets mined from recent papers. Each links to the papers that use it.
The S&P 500 is a stock market index that contains 500 of the largest publicly traded companies in the United States and is used to evaluate the overall performance of the U.S. stock market.
Electricity Transformer Temperature โ a multivariate time-series benchmark for long-horizon forecasting.
The 'GIFT-Eval' dataset/benchmark is used to evaluate the performance of time series foundation models by providing a standardized set of tasks and metrics for assessing their effectiveness in learning transferable representations across diverse temporal patterns.
ERA-5 is a reanalysis dataset that contains comprehensive atmospheric data used to evaluate and improve multi-variable weather forecasting models.
100,000 time series across domains and frequencies from the M4 forecasting competition.
DepositPhotos Traffic & Transportation Dataset (Sample) Overview This dataset is a curated sample of high-resolution Traffic, Transportation, and Urban Mobility images sourced from the DepositPhotos library. This subset is designed for training, testing, and evaluating Computer Vision models for autonomous driving, smart city infrastructure, and Generative AI applications focused on urban environments. This sample demonstrates the exact quality, diversity, andโฆ See the full description on the dataset page: https://huggingface.co/datasets/Depositphotos/traffic.
MIMIC-III is a publicly available critical care database that contains de-identified health data from patients admitted to intensive care units, and it is used to evaluate predictive models for clinical time-series analysis.
The 'PEMS' dataset is a benchmark that contains multivariate traffic series data used to evaluate traffic forecasting models, particularly in the context of extreme events and their complex spatio-temporal correlations.
The 'Electricity' dataset is a multivariate time series dataset used to evaluate forecasting methods by capturing the interdependencies among different variables over time.
The M5 dataset is a large-scale real-world retail dataset used to evaluate time series forecasting models, capturing complex dynamics across multiple variables.
The 'COVID-19' dataset/benchmark contains multi-modal data related to the pandemic, including epidemiological time series, public health policies, and demographic information, and is used to evaluate forecasting models for the short-term spread of the disease.
The M-4 competition dataset is a benchmark that contains a diverse set of time series data used to evaluate the performance of forecasting methods.
The 'Weather' dataset is a benchmark used to evaluate time series imputation methods by providing time recordings of weather-related data with missing values.
Pre-generated numpy arrays of MackeyGlass time series, generated with the jitcdde library. Please note that due to lower-level solvers used in the library, different machines, even with the same ISA and library versions, may produce different data. Thus, please use the pre-generated data included here. The dataset contains 14 time series, each uses MG parameters beta=0.2, gamma=0.1, n=10. tau is varied per time series from 17 to 30. Each time series is 50 Lyapunov times in length, with 75โฆ See the full description on the dataset page: https://huggingface.co/datasets/NeuroBench/mackey_glass.
The 'ILI' dataset contains time series data related to influenza-like illness prevalence and is used to evaluate forecasting methods for time series data.
The M-3 dataset is a benchmark for evaluating time series forecasting models, containing a diverse set of time series data across various domains.
PEMS-04 is a benchmark dataset used for evaluating spatiotemporal forecasting techniques, containing traffic data that reflects the complex spatiotemporal dynamics of transportation systems.
PEMS-08 is a benchmark dataset used for evaluating spatiotemporal forecasting techniques, containing traffic data that captures the complex spatiotemporal dynamics of transportation systems.
The 'Bitcoin' dataset contains daily close-price data used to evaluate the performance of quantile deep learning models for multi-step ahead time series prediction, particularly under high volatility and extreme conditions.
The 'Lorenz' dataset contains chaotic time series data used to evaluate the performance of time series prediction models, particularly in capturing intrinsic patterns and temporal dynamics.
The S&P 500 index is a stock market index that contains the daily and hourly closing prices of 500 large-cap U.S. companies, and it is used to evaluate short-term stock market trends and the efficacy of predictive models.
The UCR Archive is a benchmark dataset that contains a collection of time series data used to evaluate the performance of various time series classification algorithms.
Ethereum is a cryptocurrency dataset used to evaluate the performance of quantile deep learning models for multi-step ahead time series prediction, specifically focusing on daily close-price data.
The 'EUR/USD' dataset contains daily foreign exchange rates between the Euro and the US Dollar and is used to evaluate forecasting models in the context of non-stationary time series.
The 'German' dataset/benchmark contains data related to continuous intraday electricity markets in Germany and is used to evaluate forecasting models for electricity prices by analyzing the dynamics of buy and sell orders in the orderbook.
The 'Los Angeles' dataset contains hourly recordings of wind speed and is used to evaluate deep learning-based probabilistic forecasting methods for wind speed.
NASDAQ is a stock market index that contains historical stock price data used to evaluate stock price forecasting models and test the efficient-market hypothesis.
The NASDAQ-100 is a stock market index that includes 100 of the largest non-financial companies listed on the NASDAQ stock exchange, and it is used to evaluate volatility forecasting models based on high-frequency trading data.
Dataset Card for MNIST Dataset Summary The MNIST dataset consists of 70,000 28x28 black-and-white images of handwritten digits extracted from two NIST databases. There are 60,000 images in the training dataset and 10,000 images in the validation dataset, one class per digit so a total of 10 classes, with 7,000 images (6,000 train images and 1,000 test images) per class. Half of the image were drawn by Census Bureau employees and the other half by high school studentsโฆ See the full description on the dataset page: https://huggingface.co/datasets/ylecun/mnist.
ล olar is a developmental corpus of 5485 school texts (e.g., essays), written by students in Slovenian secondary schools (age 15-19) and pupils in the 7th-9th grade of primary school (13-15), with a small percentage also from the 6th grade. Part of the corpus (1516 texts) is annotated with teachers' corrections using a system of labels described in the document available at https://www.clarin.si/repository/xmlui/bitstream/handle/11356/1589/Smernice-za-oznacevanje-korpusa-Solar_V1.1.pdf (in Slovenian).
The 'Beijing' dataset contains multi-year pollutant and meteorological time-series data used to evaluate the performance of various forecasting models for hourly PM2.5 prediction in Beijing, China.
The Dow Jones Industrial Average is a stock market index that contains 30 significant publicly traded companies in the United States and is used to evaluate the overall performance of the stock market and the economy.
The 'ECG' dataset/benchmark contains electrocardiogram data used to evaluate the performance of models in capturing multivariate time series patterns and dependencies.
The 'France' dataset contains fertility data used to evaluate forecasting methods for predicting future fertility rates, specifically focusing on the uni-modal pattern of fertility with respect to age.
The Global Energy Forecasting Competition 2014 is a benchmark dataset used to evaluate multi-horizon short-term load forecasting methods in power systems.
The Japanese Mortality Database contains regional age-specific mortality rates used to evaluate the accuracy of forecasting methods for these rates at national and sub-national levels.
The 'LargeST' dataset/benchmark contains large sets of synchronized spatial-temporal time series data and is used to evaluate the performance of forecasting models in handling challenges such as the emergence and disappearance of entities.
The Lorenz-63 dataset is a low-dimensional chaotic system used to evaluate the forecasting performance of different computational models, particularly in terms of their ability to handle long-term temporal dependencies and computational efficiency.
The Lorenz system is a canonical chaotic system used to evaluate forecasting models for chaotic time series prediction.
MIMIC-IV is a publicly available dataset containing de-identified health data from critical care patients, used to evaluate predictive models for various medical outcomes.
The NIFTY-50 is a standard dataset that contains stock market data used to evaluate various forecasting models, including traditional and deep learning approaches, based on metrics such as MSE, RMSE, MAPE, POCID, and Theil's U.
PEMS-03 is a benchmark dataset used for evaluating spatiotemporal forecasting techniques, specifically in the context of transportation data.
PEMS-07 is a benchmark dataset used for evaluating spatiotemporal forecasting techniques, containing traffic data that reflects complex spatiotemporal patterns.
The 'Tourism' dataset contains time series data related to tourism metrics and is used to evaluate the performance of forecasting models in predicting future values in this domain.
WeatherBench is a benchmark dataset used to evaluate the performance of weather forecasting models, specifically focusing on short-term predictions and their accuracy in terms of root mean squared error.
fred_md (TsFile format) 107 monthly time series showing a set of macro-economic indicators from the Federal Reserve Bank. This repository contains the full source .tsf series from the Monash Time Series Forecasting Repository converted to Apache TsFile format. Summary Source dataset: Monash-University/monash_tsf Original source: https://zenodo.org/record/4654833 Monash subset: fred_md Modalities: Time-series Source series: 107 Rows: 77,896 flattened timestampedโฆ See the full description on the dataset page: https://huggingface.co/datasets/VGalaxies666/fred_md.
ercot (TsFile format) This repository contains time-series forecasting data stored in Apache TsFile format. Summary FEV subset: ercot Unified source collection: autogluon/fev_datasets Original source: https://github.com/ourownstory/neuralprophet-data/tree/main/datasets_raw/energy Series: 8 Modalities: Time-series TsFile rows (flattened observations): 1,299,648 Frequencies: 1D, 1H, 1M, 1W TsFile files: 5 Time precision: milliseconds (INT64). Licensing andโฆ See the full description on the dataset page: https://huggingface.co/datasets/VGalaxies666/ercot.
EEG data consists of electrical activity recordings from the brain, used to evaluate the effectiveness of predictive models in forecasting critical events, such as seizures.
The Australian National Electricity Market is a dataset that encompasses electricity price data across five regions, used to evaluate the performance of forecasting models in the context of high volatility and price fluctuations.
The 'Belgium' dataset contains 2024 day-ahead auction electricity prices used to evaluate the effectiveness of various pre-trained time series models for electricity price forecasting.
The Caltrans Performance Measurement System (PeMS) is a dataset that contains real-world traffic data used to evaluate the performance of various models in forecasting traffic speed and flow.
The 'Chronos-ZS' dataset/benchmark is designed to evaluate zero-shot forecasting methods for time series data by providing a diverse set of time series tasks and scenarios.
The CSI-300 is a stock market index that includes 300 of the largest and most liquid stocks listed on the Shanghai and Shenzhen stock exchanges, and it is used to evaluate stock forecasting models.
The 'CSI 500' is a dataset that contains stock market data for 500 companies listed on the Chinese stock exchanges, and it is used to evaluate forecasting methods in the context of multivariate time series analysis.
Didi Chuxing is a ride-hailing dataset used to evaluate forecasting models for spatio-temporal demand and supply-demand gaps in ride-hailing systems.
The Dow Jones Index is a benchmark that contains constituents of the Dow Jones Industrial Average and is used to evaluate the forecasting performance of Value-at-Risk (VaR) models in financial risk management.
The 'ETTh' dataset is an energy time-series benchmark used to evaluate the performance of forecasting models, particularly in the context of complex temporal dependencies and multi-source data integration.
The 'EUR-GBP' dataset contains exchange rate data between the Euro and the British Pound, and it is used to evaluate forecasting methods in non-stationary environments.
The 'German electricity market' dataset contains long-term time series data used to evaluate the accuracy of electricity price forecasting methods, specifically in the context of applying the Autoregressive Hybrid Nearest Neighbors (ARHNN) method.