Agent 如何在环境和反馈中学习
Published:
从 Reward、State、Credit 三个角度理解 Agent 训练闭环:Reward 决定学习方向,State 决定训练发生的位置,Credit 决定结果归因于哪些行动。
Published:
从 Reward、State、Credit 三个角度理解 Agent 训练闭环:Reward 决定学习方向,State 决定训练发生的位置,Credit 决定结果归因于哪些行动。
Published:
Understanding the agent training loop through Reward, State, and Credit: Reward sets the direction of learning, State sets where training happens, and Credit decides which actions the outcome is attributed to.
Published:
这个博客是一次蓄谋已久的冒险。我不止一次地和朋友们提起过我想写博客,我也不止一次地尝试开始写类似的东西。知乎、小红书、一个注册后一篇也没发过的公众号,在零零碎碎的尝试后,我又一次回到起点。
Published:
This blog is an adventure I have been plotting for a long time. More than once I told friends that I wanted to write a blog, and more than once I tried to begin something like it. Zhihu, Xiaohongshu, a public account I registered and never posted to: after all these scattered attempts, I have returned to the beginning again.