Psy-Insight is a bilingual, interpretable multi-turn dataset for mental health counseling dialogues. It includes 6,208 rounds of multi-turn counseling dialogues in English and 5,776 rounds in Chinese, annotated with step-by-step reasoning labels and multi-task labels. This dataset is designed to support the application of large language models in mental health and is suitable for tasks such as emotion classification and psychological treatment interpretation.
Psy-Insight is a comprehensive dataset designed to support the development of AI applications in mental health counseling. It includes detailed multi-turn dialogues, emotional labels, psychological treatment methods, and step-by-step reasoning annotations. This dataset is ideal for researchers and developers looking to fine-tune large language models for mental health applications.
FineWeb is a dataset of over 15 trillion tokens of cleaned and deduplicated English web data from CommonCrawl. It is optimized for LLM performance and processed using the datatrove library. The dataset aims to provide high-quality data for training large language models and outperforms other commonly used web datasets.We’re on a journey to advance and democratize artificial intelligence through open source and open science.
The American National Mental Health Services Survey (N-MHSS) is an annual survey conducted by the Substance Abuse and Mental Health Services Administration (SAMHSA) to collect data on mental health treatment facilities across the United States. The survey provides detailed information on the services and characteristics of these facilities, helping to inform policy and improve mental health care.
tartuNLP/reddit-anhedonia by huggingface-mirror (hf-mirror)