Publications

On the Role of Anticausal Direction in LLM-based Data Synthesis

Abstract

Large Language Models (LLMs) are increasingly used to generate synthetic data. Most LLM-based data synthesis workflows are anticausal : the user injects Y into the prompt to enforce targeted generation of X (Y\rightarrow X). This anticausal direction seems contradictory to the natural direction of data synthesis in machine learning and data mining: raw content X is produced first and the supervision signal Y is assigned afterward (X\rightarrow Y), i.e., the causal direction. This contrast impels us to investigate if a synthesis direction can impact synthetic data quality and downstream utility. We therefore construct paired causal and anticausal synthetic datasets for five tasks, sentiment analysis, causal language detection, ideology detection, math reasoning, and summarization. We then evaluate them via supervised fine-tuning and post-training models across these tasks. Our study finds that models trained on causal …

Date
2026
Authors
Bohan Jiang, Pingchuan Ma, Zhen Tan, Zhuoyu Shi, Fred Morstatter, Adrienne Raglin, Huan Liu
Book
Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 2
Pages
2086-2097