Publications
On the Role of Anticausal Direction in LLM-based Data Synthesis
Abstract
Large Language Models (LLMs) are increasingly used to generate synthetic data. Most LLM-based data synthesis workflows are anticausal : the user injects Y into the prompt to enforce targeted generation of X (Y\rightarrow X). This anticausal direction seems contradictory to the natural direction of data synthesis in machine learning and data mining: raw content X is produced first and the supervision signal Y is assigned afterward (X\rightarrow Y), i.e., the causal direction. This contrast impels us to investigate if a synthesis direction can impact synthetic data quality and downstream utility. We therefore construct paired causal and anticausal synthetic datasets for five tasks, sentiment analysis, causal language detection, ideology detection, math reasoning, and summarization. We then evaluate them via supervised fine-tuning and post-training models across these tasks. Our study finds that models trained on causal …
- Date
- 2026
- Authors
- Bohan Jiang, Pingchuan Ma, Zhen Tan, Zhuoyu Shi, Fred Morstatter, Adrienne Raglin, Huan Liu
- Book
- Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 2
- Pages
- 2086-2097