1000093679
Aug 08, 2026 03:09
· 2:27
· English
· Whisper Turbo
· 2 _Gözleg
Bu transkripiň möhleti geçýär 27 günler
Daşyndan gaýd etmek üçin täzele →
Diňe görkez
0:00
S…
Speaker 2 (1000093679)
Poor memory may be the prize of good generalization.
0:02
S…
Speaker 2 (1000093679)
So children having poor memory is a feature,
0:05
S…
Speaker 1 (1000093679)
not a bug.
0:06
S…
Speaker 2 (1000093679)
I want you to run with me on this for a minute.
0:07
S…
Speaker 2 (1000093679)
Models trained heavily on one domain lose what Karpathy calls entropic
0:12
S…
Speaker 1 (1000093679)
diversity.
0:12
S…
Speaker 1 (1000093679)
The wide,
0:13
S…
Speaker 2 (1000093679)
unpredictable distribution of outputs that make a system genuinely flexible.
0:18
S…
Speaker 2 (1000093679)
This is called over fitting.
0:19
S…
Speaker 2 (1000093679)
The model fits its training distribution so precisely that it stops
0:23
S…
Speaker 2 (1000093679)
generalizing beyond it.
0:24
S…
Speaker 2 (1000093679)
The output starts converging.
0:26
S…
Speaker 2 (1000093679)
The weird edge case tail end responses disappear.
0:29
S…
Speaker 2 (1000093679)
What's left is statistically central,
0:32
S…
Speaker 2 (1000093679)
confident and narrow and I've spoken about this multiple times.
0:34
S…
Speaker 2 (1000093679)
Why model's output seems very generic.
0:36
S…
Speaker 1 (1000093679)
Nothing surprising.
0:37
S…
Speaker 2 (1000093679)
Three things drive this and they compound.
0:39
S…
Speaker 2 (1000093679)
First is reinforcement learning penalizes diversity.
0:42
S…
Speaker 2 (1000093679)
So what RL does is it upgrades the output that gets rewards.
0:46
S…
Speaker 1 (1000093679)
Basically, correctness,
0:47
S…
Speaker 1 (1000093679)
conciseness,
0:48
S…
Speaker 1 (1000093679)
human approval.
0:49
S…
Speaker 2 (1000093679)
And diversity has no reward signal,
0:51
S…
Speaker 2 (1000093679)
so it gets optimized away.
0:52
S…
Speaker 2 (1000093679)
Second is high frequency pattern dominance.
0:54
S…
Speaker 2 (1000093679)
So data that appears billions of times in the training data set obviously gets
0:59
S…
Speaker 2 (1000093679)
baked deeper into the weights than the rare data.
1:01
S…
Speaker 2 (1000093679)
So the model defaults towards common pattern when sampling.
1:04
S…
Speaker 2 (1000093679)
The tails where the unusual generally generalized outputs live gets progressively
1:09
S…
Speaker 1 (1000093679)
underrepresented.
1:10
S…
Speaker 1 (1000093679)
Third,
1:10
S…
Speaker 2 (1000093679)
synthetic data accelerates both the above problems.
1:14
S…
Speaker 2 (1000093679)
when model outputs re -enter the training pool is also a real that I've made
1:18
S…
Speaker 1 (1000093679)
before.
1:18
S…
Speaker 2 (1000093679)
They amplify whatever's already dominant.
1:21
S…
Speaker 2 (1000093679)
So if everything on the internet is essentially LLM -generated output,
1:25
S…
Speaker 2 (1000093679)
then the newer models on the internet are actually progressively training on that data,
1:29
S…
Speaker 1 (1000093679)
which is synthetic data.
1:30
S…
Speaker 2 (1000093679)
Each generation of training narrows the distribution further.
1:33
S…
Speaker 2 (1000093679)
The deterioration is invisible at the sample level.
1:35
S…
Speaker 2 (1000093679)
It only shows when you look at the distribution across many outputs or when you ask
1:39
S…
Speaker 2 (1000093679)
the model to do something genuinely off its well -worn path.
1:43
S…
Speaker 2 (1000093679)
But Karpathy said something recently that I think reframes the whole problem.
1:46
S…
Speaker 2 (1000093679)
He said that humans also collapse over time.
1:48
S…
Speaker 2 (1000093679)
Children haven't overfit yet and they will always say stuff that will shock you.
1:52
S…
Speaker 2 (1000093679)
But adults end up revisiting the same thought and their learning rate goes down.
1:56
S…
Speaker 2 (1000093679)
So basically children are high entropy systems.
1:59
S…
Speaker 2 (1000093679)
They say unexpected things,
2:00
S…
Speaker 2 (1000093679)
make strange analogies,
2:01
S…
Speaker 2 (1000093679)
ask questions.
2:02
S…
Speaker 2 (1000093679)
Their mental distribution is wide.
2:04
S…
Speaker 2 (1000093679)
No strong priors yet.
2:05
S…
Speaker 2 (1000093679)
They are also terrible at recall.
2:07
S…
Speaker 2 (1000093679)
They forget almost everything from their early years.
2:09
S…
Speaker 2 (1000093679)
Adults are lower entropy.
2:10
S…
Speaker 2 (1000093679)
They pattern match rapidly and reliably.
2:12
S…
Speaker 2 (1000093679)
They also stop noticing things that don't fit their model.
2:15
S…
Speaker 2 (1000093679)
Their memory is better,
2:17
S…
Speaker 2 (1000093679)
but their flexibility is worst.
2:18
S…
Speaker 2 (1000093679)
So the question that emerges is,
2:20
S…
Speaker 2 (1000093679)
is that a coincidence or is there something about high memory capacity that structurally
2:24
S…
Speaker 2 (1000093679)
trades off against generalization?
2:25
S…
Speaker 1 (1000093679)
That's in Reeldo.
Bu transkript AI (otomat söz tanamak) tarapyndan emele getirildi. Hatalar bar bolup biler - örän möhüm ulanmak üçin ahyrky ses bilen deňle. AI düzgüni
_Dürs
Bu transkripiň AI haýalnamasyny emele etmek üçin Çap et
Çap edilýär...
Bu transkripsiýa hakda AI-den sora
Bu transkripsiýa hakda bir zat soraň — AI degişli bölümleri tapyp jogap berer.