A large-scale dataset of more than 1 million real-world user interactions with ChatGPT collected between April 9, 2023 and May 1, 2024. It captures a wide range of languages (more than 68 languages detected), user prompts, and conversational contexts. The dataset was developed by offering free access to ChatGPT and GPT-4, with participants consenting to share their chat histories for research purposes. The data includes metadata such as time stamps, hashed IP addresses (coding an IP address for privacy), country/state, and request headers.
The dataset captures a wide variety of user prompts and conversation themes, including: