[1902.04114] Using Embeddings to Correct for Unobserved Confounding
"We consider causal inference in the presence of unobserved confounding. In particular, we study the case where a proxy is available for the confounder but the proxy has non-iid structure. As one example, the link structure of a social network carries information about its members. As another, the text of a document collection carries information about their meanings. In both these settings, we show how to effectively use the proxy to do causal inference. The main idea is to reduce the causal estimation problem to a semi-supervised prediction of both the treatments and outcomes. Networks and text both admit high-quality embedding models that can be used for this semi-supervised prediction. Our method yields valid inferences under suitable (weak) conditions on the quality of the predictive model. We validate the method with experiments on a semi-synthetic social network dataset. We demonstrate the method by estimating the causal effect of properties of computer science submissions on whether they are accepted at a conference."
yesterday
The Bias Is Built In: How Administrative Records Mask Racially Biased Policing by Dean Knox, Will Lowe, Jonathan Mummolo
"Researchers often lack the necessary data to credibly estimate racial bias in policing. In particular, police administrative records lack information on civilians that police observe but do not investigate. In this paper, we show that if police racially discriminate when choosing whom to investigate, using administrative records to estimate racial bias in police behavior amounts to post-treatment conditioning, and renders many quantities of interest unidentified---even among investigated individuals---absent strong and untestable assumptions. In most cases, no set of controls can eliminate this statistical bias, the exact form of which we derive through principal stratification in a causal mediation framework. We develop a bias-correction procedure and nonparametric sharp bounds for race effects, replicate published findings, and show traditional estimation techniques can severely underestimate levels of racially biased policing or even mask discrimination entirely. We conclude by outlining a general and feasible design for future studies that is robust to this inferential snare."
