Transparency (in ML)
Concrete Problems in AI Safety — inherited
An adjacent research area concerned with understanding what complicated machine learning systems are doing internally.
Nobody fully understands what a complicated machine learning system is doing inside, and much of AI safety work quietly depends on that changing. Transparency research pursues the question directly; Concrete Problems in AI Safety (2016) names the field in a single line, declining to develop it even though several of its own proposals assume it.
What follows takes up a field named mainly to be set aside, the tool the paper's own proposals quietly presuppose anyway, and where the area sits relative to the rest.
A field named to be set aside
Transparency research, as Concrete Problems uses the term, predates the paper as an area concerned with understanding what complicated machine learning systems are doing internally, and the paper cites it as one of several related fields it declines to develop: "How can we understand what complicated ML systems are doing?" (concrete-problems, §"Transparency:", p. 21). In the broader ML literature this work is more often called interpretability or explainability, and includes techniques ranging from inspecting which input features most influence a model's output to building models whose internal structure is inherently easier for a human to reason about than a large neural network's is.
A tool the paper's own proposals often assume
Transparency's relationship to the rest of Concrete Problems is not purely peripheral, even though the paper treats it as an adjacent field: several of the paper's own remedies elsewhere implicitly assume some transparency capability rather than providing one. A Reward autoencoder (goal transparency), for instance, is only useful for creating "goal transparency" if some decoding or interpretability method already exists to extract a legible objective from an agent's actions; the paper's proposal builds on transparency as a resource rather than solving for it. This is the sense in which transparency functions in the corpus as infrastructure the accidents framework often leans on without treating its development as its own research problem.
Where it sits
Transparency-ml belongs to the theme Drawing the boundary: what accident risk is not alongside Privacy (in ML), Fairness (in ML), Security (attacks against ML systems), Abuse (of ML systems), and Policy (economic/social impacts of ML), the six fields the paper names in its closing section and explicitly declines to cover, while affirming that "fruitful intersection is likely to exist" between these topics and the paper's own five problems (concrete-problems, §"Related Problems in Safety:", p. 21).