NASA Logo

NTRS

NTRS - NASA Technical Reports Server

Press Enter or click the Search button to begin your search.

Back to Results
QUICR-learning for Multi-Agent CoordinationCoordinating multiple agents that need to perform a sequence of actions to maximize a system level reward requires solving two distinct credit assignment problems. First, credit must be assigned for an action taken at time step t that results in a reward at time step t > t. Second, credit must be assigned for the contribution of agent i to the overall system performance. The first credit assignment problem is typically addressed with temporal difference methods such as Q-learning. The second credit assignment problem is typically addressed by creating custom reward functions. To address both credit assignment problems simultaneously, we propose the "Q Updates with Immediate Counterfactual Rewards-learning" (QUICR-learning) designed to improve both the convergence properties and performance of Q-learning in large multi-agent problems. QUICR-learning is based on previous work on single-time-step counterfactual rewards described by the collectives framework. Results on a traffic congestion problem shows that QUICR-learning is significantly better than a Q-learner using collectives-based (single-time-step counterfactual) rewards. In addition QUICR-learning provides significant gains over conventional and local Q-learning. Additional results on a multi-agent grid-world problem show that the improvements due to QUICR-learning are not domain specific and can provide up to a ten fold increase in performance over existing methods.
Document ID
20060051800
Acquisition Source
Ames Research Center
Document Type
Preprint (Draft being sent to journal)
Authors
Agogino, Adrian K.
(NASA Ames Research Center Moffett Field, CA, United States)
Tumer, Kagan
(NASA Ames Research Center Moffett Field, CA, United States)
Date Acquired
September 7, 2013
Publication Date
January 1, 2006
Subject Category
Mathematical And Computer Sciences (General)
Meeting Information
Meeting: Twenty-FIrst National Conference on Artificial Intelligence
Location: Boston, MA
Country: United States
Start Date: July 16, 2006
End Date: July 20, 2006
Distribution Limits
Public
Copyright
Work of the US Gov. Public Use Permitted.
No Preview Available