Preference for explainable AI
This research investigates the demand for explainable AI in high-stakes credit settings. It finds that decision-makers strategically avoid explanations that reveal racial or gender bias to preserve moral wiggle room. Additionally, behavioural biases cause individuals to undervalue explanations even when they complement private information and improve decision accuracy.
Please login or join for free to read more.
OVERVIEW
Introduction
AI explainability, defined as the ability to explain how a prediction was made, has transitioned from a technical concern to a primary constraint in high-stakes predictive systems. This research investigates two distinct motivations for seeking explanations: accountability and human-AI collaboration (HAIC). Accountability involves auditing whether protected attributes shape decisions, while HAIC enables decision-makers to combine algorithmic signals with contextual knowledge. Modern AI facilitates the unbundling of these components, allowing high-performing scalar predictions to be consumed cheaply whilst interpretability remains a separate, optional add-on. This creates a margin of choice where one can consume a prediction while avoiding auditable meta-information about the decision technology.
Explainable AI
Explainable AI (XAI) emerged as a response to the complexity of machine learning models and their black-box nature. While early systems were transparent but lacked power, deep learning and ensemble methods required clearer insights into their outputs. Post hoc explanations for complex models focus on dissecting internals through feature importance or surrogate means. One widely used approach is Shapley-Additive-Explanation-Values (SHAP), which treats each feature as a player in a cooperative game. SHAP calculates a feature’s contribution by averaging its marginal effect across all possible subsets of features, thereby fairly distributing the total prediction among them. This provides a principled way to explain specific instances grounded in fair payoff distributions.
Conceptual overview
The study examines the demand for two separable informational objects: a prediction and an explanation. Incentives can shift demand for these objects in opposite directions, creating scope for selective transparency. The framework identifies two forces: strategic avoidance and systematic misvaluation. Strategic avoidance occurs when explanations increase accountability costs, threatening to eliminate moral wiggle room for decision-makers. Conversely, systematic misvaluation happens when the usefulness of an explanation is contingent on private context. The research predicts that increasing stakes will increase demand for AI predictions but decrease demand for explanations when protected attributes are salient.
Experimental design
In collaboration with a private U.S. lender, the study analysed real $10,000 loan allocation decisions. Online participants acted as loan officers reviewing borrower profiles that included attributes such as race, gender, employment status, and income. Participants were randomised into a neutral treatment with a fixed bonus of $0.50 or a lender-aligned treatment where a $1.00 bonus depended on loan repayment. Before deciding, participants chose whether to view the AI prediction and its accompanying explanation. A secondary experiment isolated cognitive frictions by having participants make incentivised predictions about repayment using a Becker-DeGroot-Marschak (BDM) mechanism to elicit willingness-to-pay (WTP) for explanations.
Main results
The results demonstrate a central divergence: higher stakes increase demand for predictions while reducing demand for transparency. Lender-aligned participants were significantly more likely to seek predictions but were 19.5% or 8.8 percentage points more likely to avoid explanations, particularly when the explanation was framed as an audit of sensitive features. When explanations explicitly revealed that race and gender contributed to higher default risk, participants were 7.4 percentage points or 40.6% more likely to shift toward an equal allocation of loans. This effect was not driven by shifts in beliefs about accuracy, as results remained unchanged when controlling for incentivised estimates of future performance.
Third-party sanctions were also prominent; participants imposed penalties equal to 13.3% of earnings on lender-aligned participants who accepted biased recommendations. Penalties were 51.0% higher for those who chose not to learn whether race and gender influenced the risk assessment. In the secondary experiment, WTP for explanations fell by 25.6% after participants received private information that made the explanation more useful, indicating a failure of contingent reasoning. However, when guided by a brief tutorial, WTP rose by 23.7%, consistent with rational updating.
Discussion and conclusion
The findings imply that markets and organisations may over-demand accuracy while under-consuming transparency. Because avoidance is strategic and rooted in image concerns, voluntary markets may underprovide explanations even when they correct socially costly errors. Governance should account for these behavioural dynamics, potentially through targeted mandates or auditor access. Regulations such as the EU AI Act may counteract self-serving incentives by collapsing moral wiggle room. Furthermore, firms could offer training modules or real-time decision aids to help employees interpret AI explanations. Ultimately, ensure that algorithmic systems are just and transparent requires addressing the behavioural dynamics that determine if and when explainability is actually used.