Structure learning in human sequential decision-making

Daniel E Acuna, Paul Schrater

PLoS computational biology · · Volume 6

DOI: 10.1371/journal.pcbi.1001003

How do people learn the structure of a decision problem?

Behavior that looks inefficient under a fixed model can make sense when a person is also learning how the environment works. This study connects human choices in sequential reward tasks with Bayesian models that learn both rewards and the structure that generates them.

What the study found

  • Learning the environment's structure changes the behavior expected of an optimal agent.
  • Participants' choices in the tested bandit tasks were consistent with near-optimal structure learning.

How the study works

The study combines Bayesian reinforcement-learning models with human experiments involving one-armed and two-armed bandit reward structures.

Scope and limitations

  • The evidence concerns controlled reward tasks; it does not establish optimal behavior in every real-world decision.
  • A model's assumptions about what the learner knows are part of the comparison, rather than a neutral baseline.

Abstract

Studies of sequential decision-making in humans frequently find suboptimal performance relative to an ideal actor that has perfect knowledge of the model of how rewards and events are generated in the environment. Rather than being suboptimal, we argue that the learning problem humans face is more complex, in that it also involves learning the structure of reward generation in the environment. We formulate the problem of structure learning in sequential decision tasks using Bayesian reinforcement learning, and show that learning the generative model for rewards qualitatively changes the behavior of an optimal learning agent. To test whether people exhibit structure learning, we performed experiments involving a mixture of one-armed and two-armed bandit reward models, where structure learning produces many of the qualitative behaviors deemed suboptimal in previous studies. Our results demonstrate humans can perform structure learning in a near-optimal manner.

Abstract from the original work, reproduced under its Creative Commons license. The overview above summarizes the study.

Cite this work

Daniel E Acuna, Paul Schrater (2010). Structure learning in human sequential decision-making. PLoS computational biology. https://doi.org/10.1371/journal.pcbi.1001003

Download BibTeX

View BibTeX
@article{acuna2010structure,
  title = {Structure learning in human sequential decision-making},
  author = {Acuna, Daniel E and Schrater, Paul},
  year = {2010},
  publication_date = {2010-12-02},
  journal = {PLoS computational biology},
  volume = {6},
  number = {12},
  pages = {e1001003},
  publisher = {Public Library of Science San Francisco, USA},
  doi = {10.1371/journal.pcbi.1001003},
  url = {https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1001003}
}

Overview checked September 7, 2026 against the publication record. Publication and preprint dates refer to the linked versions.

The locally hosted PDF is an unchanged copy from the original source, shared under its Creative Commons license. Copyright remains with the credited authors or rights holders.