Rational agents

Vacuum-cleaner agent

Runs the reflex vacuum agent and other agent programs in the two-square vacuum world step by step, scores them with a performance measure, and compares their average scores over all initial states.

  • Lecture reference: Rational Agents · slides 3–5
  • Lecture reference: Solving Problems by Searching · slide 8

Setup

+1 per clean square after each time step. After each time step, each clean square becomes dirty with probability p (0: dirt never comes back). The seed fixes the random dirt and the random agent's choices.

Task environment
  • Partially observable
  • Deterministic
  • Sequential
  • Static
  • Discrete
  • Single agent
  • Known
Lecture reference: Rational Agents · slides 9–16

Vacuum world t = 1 of 10

AB
Percept
[A, Dirty]
Action
Suck
Score
1 (+1)
Time step 1 of 10
t = 1: percept [A, Dirty], action Suck, score 1. Rule “if status = Dirty then return Suck” fires. Square A is cleaned. Reward +1.

Agent program Reflex agent

The Vacuum-Agent program of slide 3: Suck if the square is dirty, otherwise move to the other square.

function Vacuum-Agent([location, status]) returns an action
if status = Dirty then return Suck(this line returns the action)
else if location = A then return Right
else if location = B then return Left

As written on the slide. Lecture reference: Rational Agents · slide 3

Score over time Clean squares

051015200246810time step t

Total 18 after 10 time steps. Click the chart to show a time step.

History 10 time steps

History of the run: percept, action, reward, and score at each time step
tStatePerceptActionRewardScore
A DD[A, Dirty]Suck +11(+1)
A CD[A, Clean]Right +12(+1)
B CD[B, Dirty]Suck +24(+2)
B CC[B, Clean]Left +26(+2)
A CC[A, Clean]Right +28(+2)
B CC[B, Clean]Left +210(+2)
A CC[A, Clean]Right +212(+2)
B CC[B, Clean]Left +214(+2)
A CC[A, Clean]Right +216(+2)
B CC[B, Clean]Left +218(+2)

Expected performance Clean squares

Agent programAveragePer stepRange
Reflex agent Best Selected19.251.9318 – 20
Reflex agent with state Best 19.251.9318 – 20
Random agent 14.651.479 – 20
Table-driven agent Best 19.251.9318 – 20

Average total score over the 8 initial states, 20 seeded runs each for programs or environments with random choices, with 10 time steps, p = 0, and the measure above. Range: the lowest and highest average of a single initial state. The table-driven agent uses the table shown when it is selected.

Rational agents

For each possible percept sequence, a rational agent should select an action that is expected to maximize its performance measure, given the evidence provided by the percept sequence and the agent’s built-in knowledge.

Lecture reference: Rational Agents · slide 4

Performance measure (utility function): an objective criterion for success of an agent’s behavior. Expected utility:

EU(action) = Σoutcomes P(outcome | action) U(outcome)

Can a rational agent make mistakes?

Lecture reference: Rational Agents · slide 4
Show answer

Is this agent rational?

Lecture reference: Rational Agents · slide 5
Show answer

How many possible states? What if there are n possible locations?

Lecture reference: Solving Problems by Searching · slide 8
Show answer