Skip to content
Jambity's Blog

Lecture 2: Bellman Equation

强化学习中的数学原理 | 第二课

RL,强化学习中的数学原理6min read...

Why return is important?

return 是我们评估 policy 好不好的手段

1786507121519

discount rate 是全程记录的,我之前误以为是只有停留在原地才会有discount rate

1786507894548

Bootstrapping! 左脚踩右脚上天

1786508286193

求解这个矩阵方程就能算出

这就是贝尔曼公式(只针对这个特殊例子)

核心思想:一个 state 的 value 来自别的 state 的 value

  • : the reward obtained after taking

都是随机变量

即都可以求期望…

The step is governed bythe following probability distributions:

  • is governed by
  • is governed by
  • is governed by

接下来可以推广到多步的 trajectory

the discounted return is

也是随机变量,因为也都是随机变量

再重复一遍,也是随机变量,同样的也会有不同的

就是 的期望,叫做 state-value function or simply state value

Remarks:

  • 是 的函数,不同 触发得到的 state value 不同
  • 是 policy 的函数,不同策略得到不同 state value,这也是写成的原因,其实就是的意思
  • 代表了状态的价值,state value 越大,指能得到更多的return,policy 更好
Q: What is the relationship between return and state value?

return 是实际的结果,state value 是估计的结果(估计的来源是policy中某几步的概率)

return 是一条结果,state value 对应的是全期望公式

开始推导,there is some math.

Consider a random trajectory:

The return can be written as

then, 根据 state value 定义:

Next, calculate the two terms, respectively.

First, calculate the first term :

Note that

  • This is the mean of immediate rewards

Second, calculate the second term :

()

Note that:

  • 即马尔科夫无记忆性,知道了 就不必要知道 了

  • 倒数两行求和能换序,只要求和是收敛的都能换序

  • This is the mean of future rewards

Therefore, we have

Highlights:

  • 这就是tmd贝尔曼公式,描述了不同状态的 state value 之间的关系
  • 包含两项:第一项 immediate reward term,第二项 future reward term
  • 这个式子不是一个单独的式子,而是对所有的状态都成立,有 n 个状态就有 n 个这种式子

  • and are state values to be calculated. Bootstrapping!

  • is a given policy. Solving the equation is called policy evaluation. Evaluate 这个 policy 是好是坏

  • and represent the dynamic model. What if the model is known or unknown? 未来再搞

1786544691950

让我们把这个例子里的所有 Bellman equation 写出来

First,consider the state value of ,下面把贝尔曼公式里的各部分算出来:

  • and 就是在状态只有 这个策略, 就是,表示概率!
  • and
  • and

【强化学习的数学原理】课程:从零开始到透彻理解(完结)】 【精准空降到 11】

1786546046339

把握好 immediate reward + future reward 的含义

  • immediate reward 就是在这个状态下采取某一个 action 时得到的单个 reward
  • future reward 是执行完这个 action 后的状态 的state value
  • 上述操作遍历所有 action,注意用加权平均

从而得到方程组:

解得:

把 代入,就能算出 state value 了

那算出来了以后干什么呢?后面再讲(improve policy)

再看一个例子

1786548767421

1786548783518

可以类似的把公式手写出来

回顾一下Bellman Equation

  • 可以发现LHS有一个state value,RHS有另一个state value,所以单从这一个公式,两个未知数,我们要求出state value是不可能的。
  • 但是上面这个式子(我们可以称之为elementwise form逐元素形式)is valid for every state ,这意味着有个像这样的等式
  • 把这几个等式合起来,就得到一系列线性方程,这可以整理成 matrix-vector form

Recall that:

通过乘法分配律,前一项直接记作

  • 处于状态,按照选择action,得到的即时奖励的期望

后一项,未来价值部分:

先对,再对求和。不过因为只依赖于,我们可以交换求和顺序,然后内部的求和式是一个全概率公式,从到,有多种

为了把所有的等式写到一起,我们要给他们标上号

然后就能得到matrix-vector form:

where:

1787060495185

为什么我们要求解state value?

给定一个policy,解出他的state value,这样的过程,我们称为policy evaluation

it’s a fundamental problem in RL, it’s the foundation to find a better policy

刚才推导的 Bellman equation in matrix-vector form

The closed-form solution 解析解形式:

  • 涉及到矩阵求逆(这个矩阵是可逆的,证明暂时省略,大致是特征值都>0),要用数学工具,所以我们一般不使用这个方法,使用下面这种迭代的方法

An iterative solution 一种迭代方法:

当趋向无穷时,那个证明的是趋向的,证明思路与不动点相关。

非常重要,我们前面理解过state value后,再理解action value就会容易多

  • State value: the average return the agent can get starting from a state
  • Action value: the average return the agent can get starting from a state and taking an action

我们为什么要关注 acton value?因为我们想知道哪个action is better

Definition:

  • action value 是关于状态动作对的函数
  • depends on
  • 是指从时间t往后所有return的折扣累加

不难发现:

recall that the state value is given by:

比较一下上面两个式子,我们可以发现实际完全等于后面那块式子

(2) and (4) are the two side of the same coin:

  • (2) shows how to obtain state value from action values.
  • (4) shows how to obtain action value from state values.

接下来要更深地理解一下action value

1787555374498

1787663465927

Comments

© Jambity. All rights reserved.......