1000093679
Aug 08, 2026 03:09
· 2:27
· English
· Whisper Turbo
· 2 اسپیکر
یہ نقل ختم ہو جاتا ہے 27 دن.
دائمی محفوظہ کے لئے بہتری →
صرف دکھائی دے رہا ہے
0:00
S…
Speaker 2 (1000093679)
Poor memory may be the prize of good generalization.
0:02
S…
Speaker 2 (1000093679)
So children having poor memory is a feature,
0:05
S…
Speaker 1 (1000093679)
not a bug.
0:06
S…
Speaker 2 (1000093679)
I want you to run with me on this for a minute.
0:07
S…
Speaker 2 (1000093679)
Models trained heavily on one domain lose what Karpathy calls entropic
0:12
S…
Speaker 1 (1000093679)
diversity.
0:12
S…
Speaker 1 (1000093679)
The wide,
0:13
S…
Speaker 2 (1000093679)
unpredictable distribution of outputs that make a system genuinely flexible.
0:18
S…
Speaker 2 (1000093679)
This is called over fitting.
0:19
S…
Speaker 2 (1000093679)
The model fits its training distribution so precisely that it stops
0:23
S…
Speaker 2 (1000093679)
generalizing beyond it.
0:24
S…
Speaker 2 (1000093679)
The output starts converging.
0:26
S…
Speaker 2 (1000093679)
The weird edge case tail end responses disappear.
0:29
S…
Speaker 2 (1000093679)
What's left is statistically central,
0:32
S…
Speaker 2 (1000093679)
confident and narrow and I've spoken about this multiple times.
0:34
S…
Speaker 2 (1000093679)
Why model's output seems very generic.
0:36
S…
Speaker 1 (1000093679)
Nothing surprising.
0:37
S…
Speaker 2 (1000093679)
Three things drive this and they compound.
0:39
S…
Speaker 2 (1000093679)
First is reinforcement learning penalizes diversity.
0:42
S…
Speaker 2 (1000093679)
So what RL does is it upgrades the output that gets rewards.
0:46
S…
Speaker 1 (1000093679)
Basically, correctness,
0:47
S…
Speaker 1 (1000093679)
conciseness,
0:48
S…
Speaker 1 (1000093679)
human approval.
0:49
S…
Speaker 2 (1000093679)
And diversity has no reward signal,
0:51
S…
Speaker 2 (1000093679)
so it gets optimized away.
0:52
S…
Speaker 2 (1000093679)
Second is high frequency pattern dominance.
0:54
S…
Speaker 2 (1000093679)
So data that appears billions of times in the training data set obviously gets
0:59
S…
Speaker 2 (1000093679)
baked deeper into the weights than the rare data.
1:01
S…
Speaker 2 (1000093679)
So the model defaults towards common pattern when sampling.
1:04
S…
Speaker 2 (1000093679)
The tails where the unusual generally generalized outputs live gets progressively
1:09
S…
Speaker 1 (1000093679)
underrepresented.
1:10
S…
Speaker 1 (1000093679)
Third,
1:10
S…
Speaker 2 (1000093679)
synthetic data accelerates both the above problems.
1:14
S…
Speaker 2 (1000093679)
when model outputs re -enter the training pool is also a real that I've made
1:18
S…
Speaker 1 (1000093679)
before.
1:18
S…
Speaker 2 (1000093679)
They amplify whatever's already dominant.
1:21
S…
Speaker 2 (1000093679)
So if everything on the internet is essentially LLM -generated output,
1:25
S…
Speaker 2 (1000093679)
then the newer models on the internet are actually progressively training on that data,
1:29
S…
Speaker 1 (1000093679)
which is synthetic data.
1:30
S…
Speaker 2 (1000093679)
Each generation of training narrows the distribution further.
1:33
S…
Speaker 2 (1000093679)
The deterioration is invisible at the sample level.
1:35
S…
Speaker 2 (1000093679)
It only shows when you look at the distribution across many outputs or when you ask
1:39
S…
Speaker 2 (1000093679)
the model to do something genuinely off its well -worn path.
1:43
S…
Speaker 2 (1000093679)
But Karpathy said something recently that I think reframes the whole problem.
1:46
S…
Speaker 2 (1000093679)
He said that humans also collapse over time.
1:48
S…
Speaker 2 (1000093679)
Children haven't overfit yet and they will always say stuff that will shock you.
1:52
S…
Speaker 2 (1000093679)
But adults end up revisiting the same thought and their learning rate goes down.
1:56
S…
Speaker 2 (1000093679)
So basically children are high entropy systems.
1:59
S…
Speaker 2 (1000093679)
They say unexpected things,
2:00
S…
Speaker 2 (1000093679)
make strange analogies,
2:01
S…
Speaker 2 (1000093679)
ask questions.
2:02
S…
Speaker 2 (1000093679)
Their mental distribution is wide.
2:04
S…
Speaker 2 (1000093679)
No strong priors yet.
2:05
S…
Speaker 2 (1000093679)
They are also terrible at recall.
2:07
S…
Speaker 2 (1000093679)
They forget almost everything from their early years.
2:09
S…
Speaker 2 (1000093679)
Adults are lower entropy.
2:10
S…
Speaker 2 (1000093679)
They pattern match rapidly and reliably.
2:12
S…
Speaker 2 (1000093679)
They also stop noticing things that don't fit their model.
2:15
S…
Speaker 2 (1000093679)
Their memory is better,
2:17
S…
Speaker 2 (1000093679)
but their flexibility is worst.
2:18
S…
Speaker 2 (1000093679)
So the question that emerges is,
2:20
S…
Speaker 2 (1000093679)
is that a coincidence or is there something about high memory capacity that structurally
2:24
S…
Speaker 2 (1000093679)
trades off against generalization?
2:25
S…
Speaker 1 (1000093679)
That's in Reeldo.
یہ نقل AI (خودکار بولنے کی پہچان) سے بنائی گئی تھی. غلطیاں ہو سکتے ہیں - اہم استعمال کے لیے اصل آڈیو کے مقابلے میں جانچیں. AI پالیسي
خلاصہ
اس نقل کا AI خلاصہ پیدا کرنے کے لئے خلاصہ کرو کلک کریں.
خلاصہ...
اس نقل کے بارے ميں AI سے پوچھو
اس نقل کے بارے میں کچھ پوچھو - AI متعلقہ حصوں کو تلاش کرے گا اور جواب دے گا۔