# Lecture 10: Learning: Operant Conditioning and Observational Learning

## Introductory Psychology

---

## Learning Objectives

By the end of this lecture, students will be able to:

1. Describe Thorndike's law of effect and Skinner's contributions to operant conditioning
2. Distinguish between positive and negative reinforcement and punishment
3. Compare schedules of reinforcement and their effects on behavior
4. Explain shaping, chaining, and the role of cognitive processes in operant conditioning
5. Describe Bandura's social learning theory and observational learning

---

## Lecture Content

### I. Foundations of Operant Conditioning

Operant conditioning is a type of learning in which behavior is strengthened or weakened by its consequences. Edward Thorndike laid the groundwork in 1898 with his law of effect, derived from experiments with cats in "puzzle boxes." Thorndike observed that cats gradually learned to press a lever to escape, with their escape time decreasing over trials in a characteristic learning curve. His principle stated that behaviors followed by satisfying consequences are more likely to recur, while those followed by annoying consequences become less likely.

B.F. Skinner (1904-1990) expanded enormously on Thorndike's work, coining the term "operant conditioning" and developing the Skinner box (operant chamber) to study behavior with precise measurement. An operant, in Skinner's terminology, is any behavior that "operates" on the environment to produce consequences. Skinner rejected mentalistic explanations and focused entirely on observable behavior and environmental contingencies. His ideas extended well beyond the laboratory: works such as *Walden Two* and *Beyond Freedom and Dignity* applied behavioral principles to the design of society.

### II. Reinforcement

Reinforcement is any consequence that strengthens a behavior — that is, increases the probability that it will recur. Positive reinforcement involves adding a desirable stimulus after a behavior: praise after studying, a food pellet after a lever press, or a paycheck after work. Negative reinforcement involves removing an aversive stimulus: taking aspirin removes a headache (reinforcing the aspirin-taking behavior), or buckling a seatbelt stops an annoying buzzer. It is essential to understand that negative reinforcement is not punishment — it increases behavior by removing something unpleasant.

Primary reinforcers are innately rewarding because they satisfy biological needs — food, water, warmth, and sexual contact. Secondary (conditioned) reinforcers acquire their reinforcing power through association with primary reinforcers; money, grades, and praise are familiar examples. A token economy is a system in which tokens (secondary reinforcers) are earned for desired behaviors and later exchanged for primary reinforcers. This approach is used in schools, psychiatric hospitals, and rehabilitation programs. The Premack principle adds that a more preferred activity can be used to reinforce a less preferred one — "You can play video games after you finish your homework."

### III. Punishment

Punishment is any consequence that weakens a behavior — that is, decreases its frequency. Positive punishment involves adding an aversive stimulus after a behavior, as in spanking, receiving a traffic ticket, or being verbally reprimanded. Negative punishment involves removing a desirable stimulus, as when a teenager loses phone privileges, a child is placed in time-out, or a driver has a license revoked.

While punishment can suppress behavior, it has significant drawbacks. It suppresses behavior only temporarily and does not teach the desired alternative. It can produce fear, anxiety, and aggression, and may inadvertently model aggressive behavior — particularly when physical punishment is used. Its effects are often situation-specific, with the behavior returning when the punisher is absent, and it can damage the relationship between the punisher and the individual being punished. Punishment is more effective when it is immediate rather than delayed, consistent rather than intermittent, paired with reinforcement of an alternative desired behavior, and accompanied by a clear explanation.

<image>A 2x2 matrix illustrating the four types of operant consequences. The columns are labeled "Add stimulus" and "Remove stimulus." The rows are labeled "Increase behavior" and "Decrease behavior." Panel A (top-left): Positive reinforcement — adding a pleasant stimulus increases behavior (icon: gold star given). Panel B (top-right): Negative reinforcement — removing an unpleasant stimulus increases behavior (icon: alarm turned off). Panel C (bottom-left): Positive punishment — adding an unpleasant stimulus decreases behavior (icon: electric shock or reprimand). Panel D (bottom-right): Negative punishment — removing a pleasant stimulus decreases behavior (icon: toy taken away). Arrows indicate the direction of behavior change in each cell.</image>

### IV. Schedules of Reinforcement

Under continuous reinforcement, every desired response is reinforced. This schedule produces the fastest acquisition of new behavior but also the most rapid extinction when reinforcement is discontinued. Partial (intermittent) reinforcement, in which only some responses are reinforced, produces slower initial learning but far greater resistance to extinction — a phenomenon known as the partial reinforcement effect.

Four basic schedules of partial reinforcement have been extensively studied. A fixed-ratio (FR) schedule delivers reinforcement after a set number of responses; piecework pay is a real-world example. This schedule produces high, steady response rates with brief pauses after each reinforcement. A variable-ratio (VR) schedule delivers reinforcement after an unpredictable number of responses, as with slot machines or sales calls. It produces the highest and most consistent response rates and is the most resistant to extinction. A fixed-interval (FI) schedule reinforces the first response after a set time period has elapsed, producing a characteristic "scallop" pattern in which response rates accelerate as the end of the interval approaches — much like checking cookies in the oven. A variable-interval (VI) schedule reinforces the first response after an unpredictable time period, producing slow, steady response rates similar to checking email sporadically throughout the day.

<image>Four cumulative response graphs (cumulative records) comparing the schedules of reinforcement. Each graph has time on the x-axis and total number of responses on the y-axis. Panel A: Fixed-ratio — steep slope with brief post-reinforcement pauses (step-like pattern). Panel B: Variable-ratio — very steep, steady slope with no pauses (steepest overall). Panel C: Fixed-interval — scalloped pattern with slow responding after reinforcement that accelerates before the next interval. Panel D: Variable-interval — moderate, steady slope with no pauses. Reinforcement deliveries are marked with small tick marks on each curve. A legend indicates that steeper slopes represent higher response rates.</image>

### V. Shaping, Chaining, and Other Operant Concepts

Shaping involves reinforcing successive approximations toward a desired behavior and is used when the target behavior does not yet occur spontaneously. Training a pigeon to bowl, for instance, requires reinforcing each step that brings the animal closer to the final performance. Shaping is fundamental to animal training and to teaching new skills to children. Chaining links a series of individual behaviors into a complex sequence, with each step serving as a discriminative stimulus for the next. In backward chaining, the last step is taught first, and earlier steps are added progressively.

A discriminative stimulus signals that reinforcement is available for a particular response. A green light indicating that pecking will produce food, or an "open" sign on a store, functions as a discriminative stimulus. Much of human behavior is under stimulus control of this kind. Stimulus generalization in operant conditioning means responding to stimuli similar to the discriminative stimulus, while stimulus discrimination means responding only to the specific stimulus that signals reinforcement.

Escape learning occurs when an organism learns a behavior that terminates an ongoing aversive stimulus, while avoidance learning involves learning a behavior that prevents the aversive stimulus from occurring in the first place. Avoidance learning is notably resistant to extinction because the organism never discovers that the threat has been removed. Learned helplessness, demonstrated by Martin Seligman, occurs when organisms exposed to inescapable aversive events stop trying to escape, even when escape later becomes possible. This phenomenon has been linked to depression in humans and has been refined through the concept of explanatory style: people who attribute negative events to internal, stable, and global causes are most vulnerable to helplessness.

### VI. Cognitive Processes in Operant Conditioning

Edward Tolman challenged the purely behaviorist view that learning requires reinforcement. His research on cognitive maps showed that rats formed mental representations of mazes even without any reward. When reinforcement was later introduced, these rats demonstrated sudden improvements in performance, revealing latent learning — learning that occurs without apparent reinforcement and is demonstrated only when motivation to perform is provided.

Wolfgang Kohler's work with chimpanzees demonstrated insight learning, in which problems are solved through sudden understanding rather than gradual trial-and-error. In the famous example, Sultan the chimp stacked boxes to reach bananas hanging from the ceiling. Research on intrinsic versus extrinsic motivation has revealed the overjustification effect: providing external rewards for activities that are already intrinsically enjoyable can actually decrease intrinsic motivation. In a classic 1973 study, Mark Lepper found that children who were rewarded for drawing subsequently drew less during free time than children who had received no reward.

### VII. Observational Learning (Social Learning Theory)

Albert Bandura (1925-2021) proposed that learning frequently occurs through observation, without the observer receiving any direct reinforcement. His famous Bobo doll experiments (1961-1963) demonstrated that children who watched an adult model aggressively attacking an inflatable doll were significantly more likely to imitate that aggressive behavior. Children who saw the model rewarded (vicarious reinforcement) were most likely to imitate, while those who saw the model punished (vicarious punishment) were less likely to imitate — although, crucially, they could still reproduce the behavior when asked, revealing an important distinction between learning and performance.

Bandura identified four processes required for observational learning. The observer must pay attention to the model's behavior, retain a memory of what was observed, have the physical and mental capacity to reproduce the behavior, and possess sufficient motivation to imitate — whether from direct reinforcement, vicarious reinforcement, or self-reinforcement. Modeling influences both prosocial behavior (children learn helping, sharing, and cooperation from positive models) and antisocial behavior (exposure to aggressive models in media, peer groups, and families increases aggression). Mirror neurons, brain cells that fire both when performing and when observing an action, may support observational learning, although their precise role remains debated.

Bandura also introduced the concept of self-efficacy — one's belief in one's ability to succeed at a particular task. Self-efficacy influences whether a person will attempt a challenge, how much effort they will invest, and how they cope with setbacks. It is built through mastery experiences, vicarious experiences (watching similar others succeed), verbal persuasion, and the interpretation of one's physiological and emotional states.

<image>A diagram of Bandura's Bobo doll experiment and observational learning model. Panel A: Three experimental conditions shown as scenes — a child watches an adult model attack the Bobo doll (aggressive model condition), a child watches an adult sit quietly (non-aggressive model condition), and a control group with no model. Panel B: Results — bar graph showing levels of aggressive behavior in children across the three conditions, with the aggressive model group showing the highest aggression. Panel C: A flowchart of Bandura's four processes: Attention (eye icon) → Retention (brain icon) → Reproduction (muscle icon) → Motivation (reward icon), leading to imitated behavior.</image>

---
