Vacuum-cleaner agent
Runs the reflex vacuum agent and other agent programs in the two-square vacuum world step by step, scores them with a performance measure, and compares their average scores over all initial states.
- Lecture reference: Rational Agents · slides 3–5
- Lecture reference: Solving Problems by Searching · slide 8
Setup
+1 per clean square after each time step. After each time step, each clean square becomes dirty with probability p (0: dirt never comes back). The seed fixes the random dirt and the random agent's choices.
Vacuum world t = 1 of 10
- Percept
- [A, Dirty]
- Action
- Suck
- Score
- 1 (+1)
Agent program Reflex agent
The Vacuum-Agent program of slide 3: Suck if the square is dirty, otherwise move to the other square.
As written on the slide. Lecture reference: Rational Agents · slide 3
Score over time Clean squares
Total 18 after 10 time steps. Click the chart to show a time step.
History 10 time steps
| t | State | Percept | Action | Reward | Score |
|---|---|---|---|---|---|
| A DD | [A, Dirty] | Suck | +1 | 1(+1) | |
| A CD | [A, Clean] | Right | +1 | 2(+1) | |
| B CD | [B, Dirty] | Suck | +2 | 4(+2) | |
| B CC | [B, Clean] | Left | +2 | 6(+2) | |
| A CC | [A, Clean] | Right | +2 | 8(+2) | |
| B CC | [B, Clean] | Left | +2 | 10(+2) | |
| A CC | [A, Clean] | Right | +2 | 12(+2) | |
| B CC | [B, Clean] | Left | +2 | 14(+2) | |
| A CC | [A, Clean] | Right | +2 | 16(+2) | |
| B CC | [B, Clean] | Left | +2 | 18(+2) |
Expected performance Clean squares
| Agent program | Average | Per step | Range |
|---|---|---|---|
| Reflex agent Best Selected | 19.25 | 1.93 | 18 – 20 |
| Reflex agent with state Best | 19.25 | 1.93 | 18 – 20 |
| Random agent | 14.65 | 1.47 | 9 – 20 |
| Table-driven agent Best | 19.25 | 1.93 | 18 – 20 |
Average total score over the 8 initial states, 20 seeded runs each for programs or environments with random choices, with 10 time steps, p = 0, and the measure above. Range: the lowest and highest average of a single initial state. The table-driven agent uses the table shown when it is selected.
Rational agents
For each possible percept sequence, a rational agent should select an action that is expected to maximize its performance measure, given the evidence provided by the percept sequence and the agent’s built-in knowledge.
Performance measure (utility function): an objective criterion for success of an agent’s behavior. Expected utility:
Can a rational agent make mistakes?
Lecture reference: Rational Agents · slide 4Show answer
Is this agent rational?
Lecture reference: Rational Agents · slide 5Show answer
How many possible states? What if there are n possible locations?
Lecture reference: Solving Problems by Searching · slide 8