Task environments
Compares task environments along the seven environment types in a table like slide 17, shows their PEAS descriptions, and lists the methods from the course preview that fit each one.
- Lecture reference: Rational Agents · slides 6–8
- Lecture reference: Rational Agents · slides 9–18
Examples of different environments
| Dimension | ||||
|---|---|---|---|---|
| Observable | Fully | Fully | Partially | Partially |
| Deterministic | Deterministic | Strategic | Stochastic | Stochastic |
| Episodic | Episodic | Sequential | Sequential | Sequential |
| Static | Static | Semidynamic | Static | Dynamic |
| Discrete | Discrete | Discrete | Discrete | Continuous |
| Single agent | Single | Multi | Multi | Multi |
| Known | Not set | Not set | Not set | Not set |
Autonomous driving Slide 17
- Observable
The sensors cannot see everything around the car, or what other drivers intend.
- Deterministic
Other traffic, pedestrians, and road conditions make outcomes uncertain.
- Episodic
Each maneuver changes the situations that follow.
- Static
Traffic keeps moving while the car decides.
- Discrete
Positions, speeds, and steering angles are real-valued; time is continuous.
- Single agent
Other drivers and pedestrians act in the same environment.
- Known
Not in the slide 17 table.
PEAS of the autonomous taxi
- Performance measure
- Safe, fast, legal, comfortable trip, maximize profits
- Environment
- Roads, other traffic, pedestrians, customers
- Actuators
- Steering wheel, accelerator, brake, signal, horn
- Sensors
- Cameras, LIDAR, speedometer, GPS, odometer, engine sensors, keyboard
Autonomous taxi, as on the slide Lecture reference: Rational Agents · slide 7
Course methods
- Deterministic environments: search, constraint satisfaction, classical planning
Needs deterministic; this environment is stochastic.
- Multi-agent, strategic environments: minimax search, games Applies
Multi-agent and stochastic; the slide notes these environments can also be stochastic.
- Episodic: Bayesian networks, pattern classifiers (Stochastic environments)
Needs episodic; this environment is sequential.
- Sequential, known: Markov decision processes (Stochastic environments) Depends on Known
Stochastic and sequential; applies if it is also known.
- Sequential, unknown: reinforcement learning (Stochastic environments) Depends on Known
Stochastic and sequential; applies if it is also unknown.
Environment types
Fully observable vs. partially observable
Lecture reference: Rational Agents · slide 10Do the agent's sensors give it access to the complete state of the environment?
- For any given world state, are the values of all the variables known to the agent?
- Fully observable
- The agent's sensors give it access to the complete state of the environment: the values of all the variables are known to the agent.
- Partially observable
- The agent's sensors give it access to only part of the state: some variables' values are not known to the agent.
Pictured on the slide: simulated robot soccer seen from above vs. humanoid robots playing soccer.
Deterministic vs. stochastic
Lecture reference: Rational Agents · slide 11Is the next state of the environment completely determined by the current state and the agent’s action?
- Is the transition model deterministic (unique successor state given current state and action) or stochastic (distribution over successor states given current state and action)?
- Strategic: the environment is deterministic except for the actions of other agents
- Deterministic
- The transition model gives a unique successor state for the current state and action.
- Stochastic
- The transition model gives a distribution over successor states for the current state and action.
- Strategic
- The environment is deterministic except for the actions of other agents.
Pictured on the slide: checkers vs. backgammon, with dice.
Episodic vs. sequential
Lecture reference: Rational Agents · slide 12Is the agent’s experience divided into unconnected single decisions/actions, or is it a coherent sequence of observations and actions in which the world evolves according to the transition model?
- Episodic
- The agent’s experience is divided into atomic episodes, and the choice of action in each episode depends only on the episode itself.
- Sequential
- A coherent sequence of observations and actions in which the world evolves according to the transition model.
Pictured on the slide: a spam filter vs. Pac-Man.
Static vs. dynamic
Lecture reference: Rational Agents · slide 13Is the world changing while the agent is thinking?
- Semidynamic: the environment does not change with the passage of time, but the agent's performance score does
- Static
- The world does not change while the agent is thinking.
- Dynamic
- The world changes while the agent is thinking.
- Semidynamic
- The environment does not change with the passage of time, but the agent's performance score does.
Pictured on the slide: a Rubik’s cube vs. a cartoon cat watching a mouse carry cheese.
Discrete vs. continuous
Lecture reference: Rational Agents · slide 14Does the environment provide a fixed number of distinct percepts, actions, and environment states?
- Are the values of the state variables discrete or continuous?
- Time can also evolve in a discrete or continuous fashion
- Discrete
- A fixed number of distinct percepts, actions, and environment states; the state variables take discrete values.
- Continuous
- The state variables (and possibly time) take continuous values, so there is no fixed number of distinct states.
Pictured on the slide: a chess diagram vs. a robot arm beside a real chessboard.
Single-agent vs. multiagent
Lecture reference: Rational Agents · slide 15Is an agent operating by itself in the environment?
- Single agent
- The agent operates by itself in the environment.
- Multi-agent
- Other agents act in the environment as well.
Pictured on the slide: a rat in a maze vs. a crowd of simulated people.
Known vs. unknown
Lecture reference: Rational Agents · slide 16Are the rules of the environment (transition model and rewards associated with states) known to the agent?
- Strictly speaking, not a property of the environment, but of the agent’s state of knowledge
- Known
- The rules of the environment (transition model and rewards associated with states) are known to the agent.
- Unknown
- The agent does not know the rules of the environment (transition model and rewards) in advance.
Pictured on the slide: Monopoly vs. a room in a 3D adventure game.
Preview of the course
Deterministic environments: search, constraint satisfaction, classical planning
Can be sequential or episodic
Environments in the table:Multi-agent, strategic environments: minimax search, games
Can also be stochastic, partially observable
Environments in the table:- Stochastic environments
Episodic: Bayesian networks, pattern classifiers
No environment in the tableSequential, known: Markov decision processes
Environments in the table:Sequential, unknown: reinforcement learning
Environments in the table:
Questions
Which of the slide 17 environments fit the setting of the search lectures: fully observable, deterministic, discrete, known?
Lecture reference: Solving Problems by Searching · slides 3–4Show answer
Why is chess with a clock semidynamic rather than static?
Lecture reference: Rational Agents · slide 13Show answer
Is known vs. unknown a property of the environment?
Lecture reference: Rational Agents · slide 16Show answer
Classify poker along the seven dimensions. Which rows of the course preview apply?