Training & Behaviour
Positive Reinforcement Training: Getting Started
Updated July 29, 2026
By Morgan Reyes
The Short Answer
The alpha-dog idea comes from a 1947 study of captive wolves in a zoo enclosure roughly 10 by 20 metres, and the researcher himself noted that the captivity prevented normal social behaviour. He never used the word "alpha." The scientist who later popularised it spent over a decade publicly asking his publisher to stop printing his own book.
What replaces it is not a philosophy, it is a mechanic. Reward-based training works on timing, and the margin is tighter than most advice admits: a marker delayed by one second cut learning from 40% of dogs to 25% in a controlled study.
So the practical core of this page is four things: mark the instant, pay immediately after, keep sessions short enough that your dog stays willing, and take the treats out of the daily food ration rather than adding them to it.
The mechanics start here if you already know the history.
Where "be the alpha" actually came from
A single 1947 paper about wolves in a zoo.
Rudolph Schenkel's "Expression Studies on Wolves" was published in Behaviour, and its own subtitle is "Captivity Observations." Up to ten unrelated wolves were housed together in an enclosure of about 10 metres by 20 metres, roughly 200 square metres.
Schenkel flagged the problem himself. He wrote that the captivity conditions prohibited the normal biological course of social behaviour, and hedged his central claim: the dominance status order of wolves, he wrote, is conditioned by situations.
Two things about that paper are routinely misreported, and one of them was misreported on this page.
First, the word "alpha" does not appear in it. Schenkel wrote about "α animals." The alpha-wolf brand is a later invention.
Second, an earlier version of this guide attributed to Schenkel a "lead wolf" enforcing rank through "incessant control and repression." That quote is real and the referent is wrong: Schenkel described a top male and a top female, each policing rivals of their own sex, not one leader dominating a pack. We had trimmed a source until it matched the caricature, on a page whose entire argument is that the caricature was never his. That is corrected, and worth naming.
The man who popularised it spent a decade trying to unpublish it
L. David Mech wrote the 1970 book that put alpha wolves into popular culture, then spent thirteen summers watching wild wolves and concluded he had been wrong.
Wild packs, it turns out, are families. Breeding pairs and their offspring. In a 1999 paper in the Canadian Journal of Zoology he wrote that the alpha concept "is particularly misleading," and offered an analogy that has aged extremely well: inferring wild wolf behaviour from captive animals is "analogous to trying to draw inferences about human family dynamics by studying humans in refugee camps."
He also noted that Schenkel had considered the family-group explanation, but only in a footnote.
Two facts here need keeping separate, because the tidy version of this story is not supported. Mech's own website states that his book remained in print until 2022, and that the alpha concept is one of its outdated pieces of information, describing breeding wolves as "merely breeders, or parents." Archived captures of that page from 2019 and 2021 additionally record his numerous pleas to the publisher to stop publishing it, a sentence the live page no longer carries.
What no source establishes is that the publisher acted because he asked. An earlier version of this page said the book was taken out of print "which it was, in 2022," implying a causal link nobody has documented. He asked for years. It stopped in 2022. Both are true; the connection between them is not evidenced.
What the evidence actually shows, including its limits
The veterinary position is unambiguous. The underlying studies are more modest than the position implies, and both facts belong here.
The American Veterinary Society of Animal Behavior's 2021 statement on humane dog training says that "based on current scientific evidence, AVSAB recommends that only reward-based training methods are used for all dog training, including the treatment of behavior problems." It adds that there is no evidence aversive methods are more effective in any context, and that for aggression there are no exceptions to this standard.
Note also that a separate, older AVSAB statement addresses dominance specifically, and its scope is narrower than it is usually quoted as: dominance theory should not be used as a general guide for behaviour modification. An earlier version of this page rendered that as "training" and cited the older statement as though it were the current one.
The strongest single study is Vieira de Castro et al. in PLOS ONE, December 2020. It compared three groups across seven Portuguese training schools, assigned by video-coding rather than self-report: reward-based (42 dogs, no aversives), mixed (22 dogs), and aversive (28 dogs). Dogs in the aversive group showed stress behaviours at a scale that is hard to argue with, with lip-licking at 55.90 against 4.11 in the reward group.
And the finding this page previously left out entirely is the most important one: the effects were not confined to the training session. Dogs trained with aversive methods showed a more pessimistic bias in a subsequent cognitive-judgement task. The effect left the training room.
Now the limits, which our earlier version omitted and which matter:
- The study is quasi-experimental, and its authors state plainly that they cannot infer a true causal relationship.
- Owners volunteered, so there is selection bias in who trains where.
- It did not test effectiveness at all. It measured welfare.
- Its cortisol measurements covered only 31 of the 90 dogs and showed a negligible effect size.
Its most useful conclusion for a reader is about dose rather than equipment: the authors found the proportion of aversive stimuli mattered more than the specific tools used. Which means an earlier version of this page was wrong to gloss the aversive category as "leash corrections, prong and choke collars." The paper is not primarily about gadgets.
The other frequently-cited source, a 2017 review by Ziv, is a narrative review rather than a systematic one: three databases, English-language only, seventeen studies, a cutoff of October 2016. Its own methodological-concerns section notes small samples, unreported effect sizes, single-coder bias, and that the physical-harm evidence rests on a single case. Its conclusion is carefully worded: although positive punishment can be effective, there is no evidence it is more effective. And the author himself writes that more studies with better methodologies are needed.
Being more careful than the consensus, in the consensus's favour, is the honest position here. You will see it claimed that reward-based training is "proven far more effective." That is stronger than Ziv supports and stronger than the 2020 study tested. What the evidence supports is that aversive methods are not more effective and do measurable welfare harm. That is a sufficient argument without inflating it.
Timing, and the one number worth memorising
A marker delayed by one second cut the proportion of dogs that learned from 40% to 25%.
That comes from a 2015 doctoral study at the University of Waikato, which taught 60 dogs a novel behaviour with 20 dogs in each of three conditions:
| Condition | Dogs that learned it |
|---|---|
| Food delivered immediately, no marker | 60% |
| Marker at the instant, food one second later | 40% |
| Marker itself delayed by one second | 25% |
Three things follow, and they are more useful than any amount of philosophy.
One second is already most of the damage. So the familiar advice to reward "within about a second" is too loose. An earlier version of this page said exactly that, and the numbers above are why it changed.
That 40-versus-25 split is the actual reason a marker exists. Not tradition, not clicker culture. A marker lets you stamp the precise instant while the food arrives a moment later, which is the only way to be accurate when your hand cannot be.
You are already marking, accidentally. The same study found that in over 75% of real-world trials, owners produced a body-movement signal before any intentional feedback, averaging 0.31 seconds. Your dog is reading your shoulders whether you intend it or not.
For scale, teaching one simple novel behaviour took a bit over nine minutes on average in that work, ranging from about three minutes to eighteen. If something is taking dramatically longer, the problem is usually the timing rather than the dog.
Three rules every trainer repeats, and none of them traces anywhere
"Pea-sized treats." "The three Ds." "Five to ten minute sessions." Not one of the three traces to a credentialed source, and each has a more useful, actually-sourced version sitting right next to it.
That is not a complaint about trainers repeating them. It is a complaint that nobody seems to have checked, on a topic where the checkable version is sitting in public documents from a veterinary college and an animal-welfare accreditation body that actually audits trainers against it.
"The three Ds" of duration, distance and distraction. The mnemonic itself has no credentialed or peer-reviewed origin anywhere we could find. The three variables it names are real, and the actual increments are below: start a stay at one to two seconds, extend by a few seconds at a time, add distance a step at a time, and don't add a new variable until the dog performs the current one every single time without a food lure. Teach the variables. Skip the label.
"Five to ten minutes" as a proven optimum. Not one, in anything we could find. What is traceable: Ohio State University's Indoor Pet Initiative recommends roughly five-minute blocks totalling at least 15 minutes a day, a Virginia Tech veterinary teaching hospital independently recommends keeping sessions short for the same reason, and one hard number exists from the lab: about nine minutes on average to teach a simple new behaviour with zero distractions, above. That's credentialed guidance plus a real benchmark, not a proven number, and it's a more honest thing to hand you than one that merely sounds precise.
The rule that actually overrides the clock comes from AnimalKind, an animal-welfare accreditation programme that audits trainers against it rather than merely suggesting it: "the duration of a training session must not continue beyond a dog's willingness to participate or physical limitations." That's stricter and more useful than any timer.
"Pea-sized." Also untraceable. What the actual finding is, and why size barely matters.
How to actually start
Pick one behaviour, mark the instant it happens, pay within a second, and stop before your dog wants to.
Choose a marker. A clicker or a short word said the same way every time. AVSAB's own materials treat a verbal marker as equally valid, so the clicker is optional equipment rather than a requirement.
Pair it before you use it. Marker, then food, several times, with no behaviour required. The marker has to mean something before it can mark anything.
Reinforce every single time at first. AAHA's behaviour management guidelines put it plainly: new behaviours are learned best when rewarded every time. Only then thin out, and the guidelines are specific that "randomly and intermittently" means more often than "seldom." University extension guidance offers a first ratio: every third or fourth time, with verbal praise in between.
But never break the marker-food pairing. This is a distinction almost nobody makes and it corrects our own earlier advice. Thinning the ratio of rewarded repetitions is normal training. Marking without paying is a different thing: work reported by Cimarelli and colleagues in 2021, and summarised in the BC SPCA's review of training methods, found that clicking without food in 40% of trials produced no learning-speed advantage and a more pessimistic cognitive bias. Our earlier phrasing, "thin out the rewards," licensed exactly that. A marker always gets paid.
Raise difficulty in increments, and only from fluency. Extension guidance is usefully concrete: get the behaviour reliable before adding difficulty, start a stay at one to two seconds, extend by a few seconds at a time, add distance a step at a time, and get it right a few times in a row before progressing. A one-to-two-minute stay is a week or more of work, not an afternoon.
Set the situation up so the right answer is easy. The technical name for this in AVSAB's glossary is antecedent arrangement, and it is the single most underrated idea in dog training: change the environment so the behaviour you want is the obvious one, rather than correcting the one you don't.
Stop on willingness, not on a clock. Professional training standards require that a session not continue beyond a dog's willingness to participate, and AVSAB's position is that a dog must always feel safe and be able to opt out. Aim for short frequent sessions, roughly five minutes, several times a day. If your dog disengages, you are done.
Treats, calories, and the arithmetic nobody does
Treats should stay around 10% of daily calories, and training sessions blow straight through that if you use normal-sized treats.
Do the arithmetic. A research session might use two hundred reinforcers. A small dog's entire daily allowance is a few tens of calories. Those numbers do not fit together, and most training advice never mentions it.
Two resolutions, both from veterinary nutrition rather than from trainers:
Make them smaller, not fewer. Tufts' veterinary nutrition service points out that a dog will like a half-inch treat as much as a two-inch one. The "pea-sized" advice everyone repeats instead traces to nobody.
Take the treats out of the food, not on top of it. Measure the day's full ration in the morning and train from it. And treat the 10% ceiling as a weekly average rather than a daily hard stop, which is what makes a heavy training day survivable.
The 10% figure, made concrete: UC Davis's veterinary teaching hospital puts it plainly: if your dog eats 100 calories a day, no more than 10 should come from treats. Worked against real foods, that's about 18 kcal for 50 g of raw baby carrots, 4 kcal for a 12 g strawberry, 30 kcal for 34 g of banana.
A widely-shared bodyweight-to-treat-calorie table, published by WSAVA's Global Nutrition Committee, does not survive a second look. Extracted twice, independently, it reads 2 kg = 14 kcal, 4 kg = 24 kcal, then a flat 32 kcal for every weight from 6 kg through 40 kg, then a jump to 146 kcal at 45 kg. A flat number across a sevenfold weight range, followed by a jump, cannot be right. The 10% principle itself isn't in question here; it's corroborated independently by WSAVA's own prose, UC Davis, and Tufts. The specific table is what we won't repeat. Use the UC Davis figures above instead.
One disclosure, since this page leans on Tufts as a source: Tufts Petfoodology's authors declare funding and paid services from several major pet-food manufacturers, which we discuss on our by-product guide. The treat-size point is not a commercial claim, and you should still know who funds the people making it.
Common early mistakes
Five, all of them about the handler rather than the dog.
- Marking late. The most common and the most costly. See the numbers above.
- Marking without paying. Breaks the association you spent time building.
- Raising difficulty before fluency. If it is not reliable at one second, it will not survive three.
- Training past willingness. A dog that has checked out is learning that training ends badly.
- Repeating a cue the dog is not responding to. In real-world observation, over 44% of owner cues got no response at all. Saying it again teaches the dog that the first four times were optional.
If training is genuinely stalling rather than just starting slowly, that is a different problem with different causes, and it deserves its own page rather than a paragraph here.
FAQ
Where did the idea of the alpha dog come from?
A 1947 study of captive wolves at Basle Zoo, in an enclosure of about 10 by 20 metres, whose author noted the captivity prevented normal social behaviour and who never used the word "alpha." The detail is above.
Did the scientist behind alpha wolf theory change his mind?
Publicly and repeatedly. He called the concept "particularly misleading" in 1999 and his own site records years of pleas to his publisher to stop printing the book that popularised it. What is and isn't established.
What do veterinary behaviour bodies actually recommend?
AVSAB's 2021 statement recommends that only reward-based methods are used for all dog training, and says there are no exceptions to that standard for aggression.
How quickly does a reward have to follow the behaviour?
Faster than "about a second." A marker delayed by one second cut learning from 40% of dogs to 25%. The numbers.
How many treats is too many when training?
Around 10% of daily calories, applied as a weekly average, with the treats taken out of the measured daily ration rather than added on top. Make them smaller rather than fewer.
How long should a training session be?
Roughly five minutes, several times a day, and stop on your dog's willingness rather than on a clock. A dog must always be able to opt out.
Bottom line: the alpha model came from ten unrelated wolves in a zoo enclosure, its author noted the captivity invalidated the social observations, and the researcher who made it famous spent over a decade asking his publisher to stop printing it. What replaces it is mechanical rather than philosophical. Mark the exact instant the behaviour happens, because one second of delay cut learning from 40% of dogs to 25%. Pay every time at first, thin the ratio later, and never mark without paying. Raise difficulty only from fluency and in small increments. Stop when your dog stops being willing rather than when a timer says so. And take the treats out of the daily ration rather than adding them on top, making them smaller rather than fewer.
Keep reading
- What Actually Counts as a "By-Product" in Pet Food. What is actually in the food you are training with.
- Read Any Pet Food Label in 60 Seconds. Including how to find the calorie figure the 10% rule depends on.
- How We Research and Fact-Check Every Guide. Including the misquote on this page and how it was found.
Three follow-ups are planned and not built: enrichment and puzzle feeders, which is about food as occupation rather than as a reinforcer; what to do when reward-based training stalls after you have competence; and separation anxiety, which is a welfare problem rather than a training one.
A Note on Scope
This guide covers basic training mechanics for a healthy dog. It is not veterinary or behavioural advice, and no veterinary behaviourist reviewed it.
Aggression, severe fear, and separation distress are not training problems and this page will not help with them. They need a veterinary behaviourist or a veterinarian, partly because a sudden behaviour change can be pain or illness presenting as temperament. If your dog has bitten, is threatening to bite, or panics when left alone, get professional help rather than a technique from the internet. A sudden unexplained change in a previously settled dog warrants a veterinary examination first.
Sources
Image credits
- “Close-up of a woman giving a treat to a dog outdoors on a sunny day.,” Liudmyla Shalimova / Pexels, Pexels License
- Mech LD, "Alpha status, dominance, and division of labor in wolf packs," Canadian Journal of Zoology 77:1196-1203, 1999. The retraction, the refugee-camp analogy, and the note that Schenkel raised the family hypothesis only in a footnote.
- Schenkel R, "Expression Studies on Wolves: Captivity Observations," Behaviour 1:81-129, 1947. The enclosure dimensions, Schenkel's own caveat about captivity, and the absence of the word "alpha." Read in English translation via an archived scan.
- AVSAB Position Statement on Humane Dog Training, 2021. The reward-only recommendation, the no-evidence-of-greater-effectiveness finding, and the no-exceptions-for-aggression standard. AVSAB's separate dominance statement is older and narrower in scope.
- Vieira de Castro AC et al., "Does training method matter? Evidence for the negative impact of aversive-based methods on companion dog welfare," PLOS ONE 15(12):e0225023, 2020. Three groups, the stress-behaviour findings, the cognitive-bias result, and the authors' own limitations.
- Ziv G, "The effects of using aversive training methods in dogs: A review," Journal of Veterinary Behavior 19:50-60, 2017. DOI 10.1016/j.jveb.2017.02.004. A narrative review, with its own methodological-concerns section.
- Browne CM, doctoral thesis, University of Waikato, 2015. The marker-timing conditions and learning rates, the acquisition times, and the accidental-body-signal finding.
- Makowska IJ and Cavalli CM, Review of dog training methods, BC SPCA, 2023. Cited by AVSAB's own statement, and the route by which we read the Cimarelli marker-pairing result.
- Tufts Petfoodology on treats and calories. Treat size and the weekly-average approach. Authors disclose pet-food industry funding.
- Ohio State University College of Veterinary Medicine, Indoor Pet Initiative, "Basic Manners" (Puppy Prep School). The 5-minute-blocks/15-minutes-a-day session guidance, the actual difficulty-raising increments, and the source that "the three Ds" itself does not trace to.
- AnimalKind Dog Training Standards, v1.8, BC SPCA accreditation programme. Standard 10.2, the session-duration rule trainers are actually audited against.
- UC Davis School of Veterinary Medicine, Nutrition Support Service, "Treat guidelines for dogs" (2020). The 100-calorie worked example and the real-food kcal figures.
- WSAVA Global Nutrition Committee, "Feeding treats to your dog" (2025). The 10% principle, corroborated independently of Tufts and UC Davis. Its own bodyweight-to-calorie table is not reproduced here; see "What we didn't use" below for why.

