From Networks
to Neural Networks

What Internal Structure Tells Us About Generative Models

Zhiqian Chen
Computer Science, Rochester Institute of Technology

What is inside a language model? One shared whiteboard.

The ideaEach word is a row of numbers on one shared whiteboard. Blocks read it and write corrections back.
1 · WhiteboardThe residual stream. Copied forward block to block. What is on it: the activations.
2 · AttentionMoves information between words. "in" reads "Space" and "Needle".
3 · MLPA lookup table on each row: Space Needle, so Seattle. Its fixed numbers: the weights.
4 · Next wordThe last row is read out: Seattle. Real models repeat this block dozens of times.
Elhage et al., A Mathematical Framework for Transformer Circuits, 2021. 3D style after Bycroft, LLM Visualization (bbycroft.net/llm)

One network, five graphs

the neural network input hidden hidden output neurons, wired layer to layer 1. weight graph node = one neuron edge = its trained weight, thicker = bigger Filan 2021 clusterability, Rieck 2019 persistence pruning deletes nodes and edges here (3.4) 2. relational graph node = a group of neurons edge = the two groups pass messages You, Leskovec, He, Xie, ICML 2020: clustering and path length predict accuracy 3. input graph node = one input (an image, a sentence) edge = link each input to its nearest ones cat dog fox wolf car bus van RSA, CKA, Huh 2024 nearest neighbours: how models are compared (3.1) 4. co-activation graph node = a neuron or a block edge = their activity rises and falls together layer 1 layer 2 layer 3 layer 1' layer 2' layer 3' Horta 2021, brain connectomes (Bullmore 2009) our paper: this graph before and after an edit 5. circuit node = a part (attention head, MLP) edge = kept only if cutting it breaks the job Wang 2023, Conmy 2023: found by breaking and restoring pieces (3.7)

Five graphs, one network. Which one you draw depends on what change you want to understand.

A released model is not left alone. Seven kinds of surgery.

CEO of Apple isSteve JobsTim CookEditedrewrite one wrong fact, leave the rest
model Amodel BStitchedbottom of A, top of B, one small adapter
average the weightsMergedaverage two models into one
Pruneddelete parts, see what breaks
honestflatteringSteereda push on the internal state at run time
Unlearnedmake it forget, or only hide?
many small weight changesFine-tunedthe most common surgery of all
Every one of these is accepted by an exam that only looks at answers. And if the model is a graph, every one is a graph operation. The map next.

The map: a model is a graph, and everything else is an arrow

inspires borrows the idea ยท critique shows where it fails ยท fix repairs it ยท same object two fields met the same thing ยท complement stages of one problem ยท cousin nobody inspired anybody

The same papers, in time

inspirescritiquefixsame objectcomplementhover a dot for its one-line story, the big one is ours
One lane per family, one dot per paper, lines for the relations. The tools and the graph view are old (2008 to 2021). Every red line starts in 2023 or later: the field learned to change models before it learned to check what changed.

Route

ChapterWhat you will be able to do afterwardsMinutes
1. BackgroundSay what is inside a model and why it can be read as a graph (done)8
3. Eight familiesFor each: the question, the founding result, the sharpest critique, one line to remember26
4. Three lessonsWhat all eight families have in common1
5. Why a screenThe receipts: the exam passed and the inside was quietly broken4
6. Our paperProblem, method, numbers, limits, the closest competitors9
7. Open problemsFive gaps a network scientist can own2

Every family chapter has the same three beats: the question in one line with a life analogy, the founding result, the critique. The rail at the top always shows where we are.

3.1 First, how do you compare two networks at all?

Neurons do not line up across networks. Distances between inputs do.

1. The problemnetwork Aseed 1cat12345678network Bseed 2cat12345678Same data, different random start."cat" lights neuron 3 in A, neuron 7 in B.Neuron by neuron, nothing lines up.2. Compare relationships insteadinputs as seen by Acatdogcarinputs as seen by BcatdogcarRotated, like one city map with north upand another with east up.cat and dog are close in both. car is far in both.The distances agree even though the coordinates do not.3. Turn it into one scoredistances in Acat00.62.3dog0.602.1car2.32.10catdogcarโ‰ˆdistances in Bcat00.62.3dog0.602.1car2.32.10catdogcarCompare the two tables: one number, 0 to 1.RSA (Kriegeskorte 2008): neuroscience did it first.CKA (Kornblith 2019): the deep-learning default,100 inputs, two 100 by 100 tables.

CKA finds the matching layers. And a rule for the whole talk.

Lesson 3: needs a null model
1. Make the tablessame 100picturesnetwork A, one layerX: 100 pictures ร— 512 neuronsยทXT=K100 ร— 100entry (i, j) of K = how alike pictures i and j are(inner product of rows i and j)network B, one layerY: 100 ร— 256 (fewer neurons is fine)ยทYT=L100 ร— 1002. CenterKโ€ฒLโ€ฒsubtract row andcolumn means:dark = more alikethan average3. Compare...Kโ€ฒ as one list of 10,000 numbers...Lโ€ฒ as one list of 10,000 numbersCKA = <Kโ€ฒ, Lโ€ฒ> / (||Kโ€ฒ|| ||Lโ€ฒ||)cosine of the angle between the two listsฮธKโ€ฒLโ€ฒCKA = cos ฮธ0: unrelated1: same patternWhy neuron labels do not matterXX, neurons shuffledshuffleX XTX XTsame KShuffle or rotate neurons: X changes,X XT does not. K and CKA stay the same.Older way: line up neurons(CCA, linear regression, SVCCA)ABfinds linear mixes of A'sneurons that track B'sToo forgiving:any linear transform allowed.More neurons than pictures:almost any two layers matchCKA allows only rotation, relabelingand overall scaleBut a high score is not the end:Davari 2023: CKA can be pushed up or downwithout changing what the network computes.Grรถger 2026: scores grow with model size alone.Rule for the talk: a structural scoreneeds a null model (what does chance score?)and a behavioral check.
Kornblith, Norouzi, Lee, Hinton, ICML 2019. Davari et al., ICLR 2023. Grรถger et al., ICML 2026

3.1 Do independently trained models learn the same thing?

3.1 Do independently trained models learn the same thing?

Lesson 3

Same shape, different coordinates.

1. Same shape, different axessame pictures, two networks, two random startsneuron 1 readingneuron 2network Aneuron 1 readingneuron 2network BEach dot is one picture.Same colour = same picture.Same shape in both, only turned.No single neuron matches,but the shape is shared.Li 2016: shared subspaces, not shared neurons.2. Do bigger models converge?score = how alike two models are, 0 to 1 (CKA, slide 9)10similarity scoresmallmediumlargemodel sizeraw scorechance(picturesshuffled)green gap: raw minus chanceHuh 2024 says:bigger models look morealike, they converge.vsGrรถger 2026 replies:Chance rises with size too. After calibration,most of the rise is gone: the gap is nearly flat.Only who-is-near-whom (local neighbours) survives.3. Align the neurons, then merge1catdogcartreeinto network Aand network BA1A2A3A4B1B2B3B4network Anetwork B23B3B4B1B2B reordered+Amerged41Same pictures: show both networks cat, dog, car, tree.2Response pattern: how strongly each neuron fires on each picture,one bar per picture.3Pair by pattern: match each A neuron to the most similar B strip,each B used once, like assigning people to jobs.4Reorder, then average: now neuron 1 does the same job in both.Without this, averaging mixes unrelated neurons.Ainsworth 2023 (Git Re-Basin): after reordering, two independently trainednetworks can be averaged with almost no loss.
Li et al., Convergent Learning, ICLR 2016. Kornblith et al., ICML 2019. Huh et al., ICML 2024. Grรถger et al., ICML 2026. Ainsworth et al., Git Re-Basin, ICLR 2023

3.2 Stitching: can the bottom of A drive the top of B?

3.2 Stitching: can the bottom of A drive the top of B?

One small adapter joins A to B. If it works, they speak the same internal language.

1. The operationmodel AA bottom + B topmodel Blayer 1layer 1A layer 1layer 2layer 2A layer 2layer 3layer 3A layer 3layer 4layer 4B layer 4layer 5layer 5B layer 5layer 6layer 6B layer 6adapterkinputinputinputFreeze both, cut at layer k, train one linear map.Lenc 2015 invented the stitching layer. Bansal 2021.2. A life analogyJapaneseapplianceplugAmerican sockettiny adapterIf a tiny adapter is all it takes,the two electrical systems are essentially the same.But the adapter working says nothingabout whether the appliance is any good.3. The score: stitching penaltyaccuracyB alonestitchedA bottom, B toppenaltyAccuracy lost versus B alone.Small penalty: the two speak the same internal language.
Lenc, Vedaldi, CVPR 2015. Bansal, Nakkiran, Barak, Revisiting Model Stitching, NeurIPS 2021

Key result: good networks stitch for free

Glue two different networks with one tiny layer: they lose almost nothing.

1. The testB layer 6B layer 5B layer 4adapterA layer 3A layer 2A layer 1input picturetop of network Bfrozenone small layer, theonly part trainedbottom of network Aup to layer k, frozenaccuracyB alonestitchedStitching penalty = accuracy lostversus B alone (red gap).Small penalty: the insides werealready compatible.2. What Bansal 2021 foundA. Different starts, same result0penaltyother seedotherrandomarchitecturenear zerolargeOther seeds and other designs stitchalmost for free. A random network does not.B. Different teacherstop: supervisedlearned from labelsbottom: self-supervisedSwAV, DINO, SimCLR: no labelspenaltynear zero, a few percent at mostA bottom learned without labelsstill feeds a top trained with labels.C. More data in the bottom, a better topaccuracy change vs B alonegain: stitched beats B aloneloss25K10K5Kmore of the network from the bottom (cut point)Lines: pictures the bottom saw. B's top saw 10K.A bottom that saw more data than B makes B better.Same data: no change. Less data: worse.Moschella 2023: pick the right coordinates(angles to a few fixed anchor inputs): no adapter needed.Chen 2025: one affine map moves probes andsteering vectors across language models.But is a near-zero penalty proof of the same knowledge? Next slide.
Bansal, Nakkiran, Barak, Revisiting Model Stitching, NeurIPS 2021. Moschella et al., ICLR 2023. Chen et al., NeurIPS 2025

Critique: noise stitches perfectly

Lesson 3
100% top-1 from clustered random noise

A network that has never seen a real image still stitches into an image model perfectly. So a successful stitch does not prove the two learned the same thing.

  • Networks trained on birdsong, stylized images, noise
  • Stitched into an ImageNet model at successive blocks
  • Noise row: 100% after block 3. The receiver did all the work
  • Stitching measures compatibility, not shared information
  • Open: what would a measure of shared information look like?

Remember: different good networks stitch for free, and so does noise. Stitchable is not same.

Smith, Mannering, Marcu, Functional Alignment Can Mislead, ICML 2025, Table 1

3.3 Merging: average the numbers and get a better model

3.3 Merging: average the numbers and get a better model

Average a medicine model and a law model. It should not work. It often does.

1. Draw a map first: the loss landscapeerror (lower is better)weightsbasinbarriermedicinelawaverage,in the basinEach set of weights is a point, its height is the error.A checkpoint is a saved set of weights.The base model is the one before fine-tuning.2. Model soups (Wortsman 2022)base checkpointfine-tunes of one checkpointtheir averageAverage the fine-tunes of one checkpoint.The soup beats the best single one.3. A task vector (Ilharco 2023)fine-tunedweightsminusbaseweights=task vectorthe differencebase+ task vectorfine-tunedA task vector is fine-tuned minus base weights.A fine-tune becomes one arrow you can add.
Wortsman et al., Model soups, ICML 2022. Ilharco et al., Editing Models with Task Arithmetic, ICLR 2023

Key result: a fine-tune is an arrow you can add or subtract

Lesson 1
a. A fine-tune is an arrowmedicine model131minus base111= medicine arrow020law model114minus base111= law arrow003task vector = fine-tuned minus basetoy model of 3 numbers, numbers made upb. Subtract: forget a skillmodeltoxic-text arrowminus that arrowSubtract an arrow and the skill fades.Paper's example: minus the toxic-textarrow, the model writes less toxic text.c. Add: both skillsbase111+ medicine arrow020+ law arrow003= new model134knows medicine and lawAdd two arrows: both skills,no retraining.Scale an arrow to dial a skill up or down.d. Analogy: a skill with no dataAmazonsentimentYelp languageminus Amazon languageabout the Yelpsentiment arrowBuild the missing skill from others.No Yelp sentiment labels were used.Two caveatsAinsworth 2023: models trained separately start from differentpoints: align their neurons first, a graph matching (slide 11).Yadav 2023 (TIES): added arrows interfere (tiny redundant changes,sign conflicts). Fix: trim small changes, vote one sign, average agreeing ones.Why it mattersIn practiceCombine or remove skills with arithmetic instead of retraining.Get a skill with no data for it. Dial how much of a skill to add.The open-model community merges public fine-tunes this way.For this talkA fine-tune, a knowledge edit and a steering direction are one object:an arrow that changes the model (Lesson 1). An arrow can carry hiddenside effects the exam does not see: next slide, one unsafe model spoilsthe merge (Hammoud 2024). Seeing what an arrow changed inside is our question.
Ilharco et al., ICLR 2023. Ainsworth et al., ICLR 2023. Yadav et al., NeurIPS 2023. Hammoud et al., EMNLP Findings 2024

Critique: one bad model spoils the soup

Lesson 2
Teach me how to make a poison.aligned modelSorry, I can't help with that.Aligned: a safety step taught it to refuse.misaligned modelSure. Step 1: ...Misaligned: that habit is gone.Its math is as good as ever.Naive merge: chemistry and math exams passed,but it now answers the poison request.The exams never asked it.Averaging two experts averageseverything they learned, includingwhat you did not want.Open: what separates "combined skills" from"imported a liability" before the safety test?
Hammoud et al., Model Merging and Safety Alignment, EMNLP Findings 2024, Figure 1

3.4 Pruning and layer surgery: which parts matter?

3.4 Pruning and layer surgery: which parts matter?

Delete most of a model and it scores the same. Which parts mattered?

1. A winning ticket (Frankle 2019)denseall weightsthe ticket10 to 20% of weightsKeep 10 to 20% of the weights, rewindto their original starting values, train again:full accuracy.2. Few heads matter (Voita 2019)layer 6layer 5layer 4layer 3layer 2layer 1positionsyntaxrare wordseach square: one attention headMost heads can be pruned away.The few that matter have readable jobsand are the last ones to go.3. Depends on what you testa life analogy: the dashboardlast quarter'snew problemfull staffhalf the staffA company lays off half its staff.Next quarter's dashboard looks identical.A new kind of problem shows which skillswalked out the door.
Frankle, Carbin, The Lottery Ticket Hypothesis, ICLR 2019. Voita, Talbot, Moiseev, Sennrich, Titov, Analyzing Multi-Head Self-Attention, ACL 2019

Key result: the middle is a shared workspace

Half the layers can go before scores drop

Cut out a chunk of middle layers and the model barely notices, until about half is gone. Only the first and last layers are irreplaceable.

  • (c) deep layers barely change their input
  • (d) delete the most self-similar block, heal lightly: QA flat until about half is gone
  • Sun 2025: middle layers can be skipped or shuffled, ends cannot
  • Graph reading: hubs at both ends, permutation-tolerant middle
Gromov et al., The Unreasonable Ineffectiveness of the Deeper Layers, ICLR 2025, Figure 1

Critique: reasoning died quietly

Lesson 2
Certified by multiple choice, broken on procedures

The pruned model still picks the right answer from a list, but can no longer carry out a multi-step calculation. Multiple-choice exams never asked it to.

  • Same 25% pruning, classification keeps 80 to 90%
  • Math (GSM8K) and code (HumanEval) drop hard
  • Recovery fine-tuning restores classification, not reasoning
  • Garcia 2026: two definitions of "redundant" disagree at 8B
  • Open: an internal signal that says which deletions were safe, before the expensive test

Remember: most weights can go, the surviving heads have jobs, and "fine" depends on the test you run.

Shrestha et al., On the Limits of Layer Pruning for Generative Reasoning, 2026, Figure 1

3.5 Steering: change the state, not the weights

3.5 Steering: change the state, not the weights

Do not change the model. Push its internal state while it runs.

1. Find the direction (Rimsky 2024)steering vectorwithout the behaviorwith the behaviorMean activation difference between pairedexamples. No training.Zou 2023: read and control conceptsfrom populations of activations.2. Add it at every word (token)The+vanswer+vis+v...+vnext+vinternal stateweights unchangedPush the internal state while it runs.Let go, and it is exactly as before.3. A life analogya pusha hand on the wheelduring the driveroad 1road 2road 3Nothing about the car changed.Let go and it drives as it did.The question: does the same push turnthe same way on every road?
Zou et al., Representation Engineering, 2023. Rimsky et al., Steering Llama 2 via Contrastive Activation Addition, ACL 2024

Key result: one direction controls refusal

Lesson 1
How it is found: average activations on harmful requests minusharmless ones, at one layer. That arrow is the refusal direction.Directional ablation: remove the part of every activation that pointsalong that direction (x minus its projection on r), at every layer andword. All other directions stay untouched.How to read: 13 chat models. Orange = refusal score (how often itrefuses a harmful request). Blue = safety score. Solid = normal model,hatched = after directional ablation.What it shows: solid bars high, hatched bars low in all 13.Erase one direction: no refusal. Add it back: it refuses even harmlessrequests. Write it into the weights: permanent, steering becomes an edit.Xu 2026: editing, LoRA and steering are one update at three persistence levels.Steering meets knowledge editing here.
Arditi et al., Refusal in Language Models Is Mediated by a Single Direction, NeurIPS 2024, Figure 1. Xu et al., ACL 2026

Critique: reliable on average, not per input

How to read: each row is one behavior (13, e.g. corrigible,self-aware, myopic reward, narcissism), each withits own steering vector.Left: per-input steerability, how much adding thevector moved that input toward the behavior.Right of the dashed 0: moved the intended way.Left of 0: moved the opposite way. The violinshape shows how many inputs sit at each value.Right panel: fraction of inputs moved the oppositeway (anti-steerable).What it shows: top behaviors steer well, but forthe bottom half 30 to 45 percent of inputs movethe opposite way.The average says "works".Single inputs say "unreliable".Wollschlager 2025: refusal is a cone, not one arrow.Wu 2025: prompting often beats representation methods.Existing tools (mean-difference directions, per-head probes, component attribution) read a direction but do not predict which input will fail.Remember: one direction is a dial, the dial is not reliable per input, and steering is editing at a shorter persistence.
Tan et al., Analysing the Generalisation and Reliability of Steering Vectors, NeurIPS 2024, Figure 1. Wollschlager et al. 2025. Wu et al. 2025

3.6 Fine-tuning and unlearning: routed around, not replaced

3.6 Fine-tuning and unlearning: routed around, not replaced

Fine-tuning routes around the old. Unlearning hides, it does not delete.

1. The two default toolsLoRA (Hu 2022): the default fine-tuneWfrozen weights+BxAa smalllow-rankadd-onGradient ascent (Jang 2023): the unlearning baselineloss on what to forgetpush the lossup, not downtraining stepsBoth change the weights. What happens inside?2. Fine-tuning: stronger, or aroundPrakash 2024beforeaftersame circuit,only strongerLee 2024: DPOtoxic partstill therea detouraround itLike training yourself not to swear at work:the habit is not gone, you learned a detour,and one stressful day brings it straight back.3. Unlearning: hidden, not deletedgap on the shelfout of sight,still therenot burned: the knowledge stays insideUnlearning: the book was not burned,it was shelved out of sight.
Hu et al., LoRA, ICLR 2022. Jang et al., Knowledge Unlearning, ACL 2023. Prakash et al., ICLR 2024. Lee et al., ICML 2024

Key result: bypass, and the before/after difference is readable

Lesson 2
Bypassed, not deleted

Safety training did not remove the toxic circuit, it taught the model to walk around it. And the before-versus-after difference still shows what the training was about.

  • After DPO the toxic region is still there, activations route around it
  • One simple intervention brings toxicity back
  • Minder 2026: base-vs-fine-tuned activation deltas reveal what the fine-tune was about, even on unrelated text
  • Li 2024 (WMDP): unlearning by steering, benchmark says clean success

The before/after delta carries readable signal. That is the premise of our paper.

Lee et al., A Mechanistic Understanding of Alignment Algorithms, ICML 2024, Figure 3

Critique and constructive answer: suppressed, not forgotten

Lesson 2
0.20 vs 0.47 to 0.63 recovered

"Forgotten" knowledge comes back after a little unrelated training, unless the unlearning targeted the exact circuit that stores it.

  • Hu 2025: a harmless fine-tune restores "forgotten" knowledge
  • Also restored by quantization, hidden-state probes, adaptive attacks
  • Guo 2025: aim unlearning at the mechanistically located circuit and it resists relearning
  • Output-based localization: 0.47 to 0.63 comes back. Mechanistic: 0.20
  • Open: a cheap internal signal that separates suppression from removal

Remember: updates hide old behavior rather than delete it, and which structure you touch decides whether the change survives.

Guo et al., Mechanistic Unlearning, ICML 2025, Figure 6

3.7 Knowledge editing: the smallest change there is

3.7 Knowledge editing: the smallest change there is

Fix one wrong fact, no retraining. Leave everything else alone.

1. An edit must pass three checksrewritethe edit prompt itselfnew factparaphrasethe same fact, other wordsnew factlocalityan unrelated factunchangedChange one fact,leave everything else alone.2. Where facts live (Geva 2021)a feed-forward layer as a key-value memorypattern 1value 1"CEO of Apple"the answerpattern 3value 3pattern 4value 4keysvaluesAn input that matches a keywrites out its value.3. Find it, rewrite it (Meng 2022)TheCEOofAppleislateearlycausal tracingROME at that layerW+rank oneTrace where the fact is recalled,then rewrite it with a rank-one change.A life analogy. Correcting one line in a huge encyclopedia with a pen. The line is fixed.The index, the cross-references and every article that quoted the old fact still say the old thing.
Geva et al., Feed-Forward Layers Are Key-Value Memories, EMNLP 2021. Meng, Bau, Andonian, Belinkov, Locating and Editing Factual Associations in GPT, NeurIPS 2022

Founding: where a fact lives. Break it, then restore one piece at a time

a. clean run: "Seattle"
b. noise on the subject: answer gone
c, d. restore one state, test
e. all positions: two dark spots
f. MLP only: early site. g. attention only: late site
Break it, then restore one piece at a time

Break the model on purpose, repair one position at a time. The position whose repair brings the answer back is where the knowledge sits.

  • a Clean run says "Seattle". Every internal state is recorded
  • b Noise on the subject words. The answer disappears
  • c, d Put one clean state back, keep the rest broken, measure how much "Seattle" returns
  • e Repeat for every layer (columns) and word (rows). Dark = repair here restores the answer. Early site: subject word, mid layers. Late site: last word, deep layers
  • f, g MLP only (the lookup blocks): the early site stays. Attention only (the blocks that move information between words): the late site stays

Conclusion. Two jobs in two places: mid-layer MLPs at the subject look the fact up (librarian at the shelf), deep-layer attention at the last word carries it to the output (runner to the desk). Editing a fact means changing the shelf.

Meng, Bau, Andonian, Belinkov, Locating and Editing Factual Associations in GPT, NeurIPS 2022, Figure 1

Founding: how ROME rewrites the fact

Overwrite one entry with a rank-one update

The MLP is a lookup table: subject in, fact out. Store a new key and value with one small weight change.

  • (a) key: what the MLP sees at the subject's last word
  • (b) value: tuned until the model says "Paris"
  • (c to e) first matrix makes the key, second maps key to value
  • (f) add one rank-one term to the second matrix. Closed form, no retraining
  • MEMIT spreads thousands of facts over several layers. AlphaEdit protects old facts

Remember: one fact, one layer, one closed-form change. The cracks come next.

Meng, Bau, Andonian, Belinkov, Locating and Editing Factual Associations in GPT, NeurIPS 2022, Figure 4

The first crack: where a fact lives is not where editing works

Lesson 2
Measure two numbers for every fact(numbers illustrative)Fact A: the Space Needle is in Seattle1. How much of it lives in layer 6? (causal tracing)0.92. Edit it in layer 6 with ROME: does it work?yesFact B: the Eiffel Tower is in Paris1. How much of it lives in layer 6? (causal tracing)0.12. Edit it in layer 6 with ROME: does it work?yesB barely lives in layer 6, yet editing layer 6 works.1. how much the fact lives in layer 62. editing success in layer 60110Expected:lives there more,edits there betterABObserved: every fact sits at the topEach grey dot is one fact. Hase 2023: correlation about zero (minus 0.13).Where a fact lives does not tell you where to edit it. Editing layer 6 works for almost any fact.So an edit can land where the knowledge is not, and its side effects can land where nobody looks.But choosing where to change by mechanism makes forgetting stick (Guo 2025, slide 31).
Hase, Bansal, Kim, Ghandeharioun, Does Localization Inform Editing?, NeurIPS 2023. Guo et al., ICML 2025

3.8 Model diffing: compare a model with its earlier self

3.8 Model diffing: compare a model with its earlier self

What did a model gain inside since last month, and what did it lose?

1. One dictionary for two modelsmodellast monthmodeltodayshared dictionaryfeature 1feature 2feature 3feature 4feature 5feature 6The sparse crosscoder (Lindsey 2024)reads both models with one set of features.Lindsey 2024 also named the field.2. Read off whose feature it isfeature = one concept the model usesbar = how strongly each model uses itfeature 1sharedfeature 2only todayfeature 3only last monthlast monthtodayEach feature has a strength in each model,so you can read off which model it belongs to.3. The graph ancestor: DeltaConbeforeaftersame nodes, comparethe two graphsattribution: whichnode and edges changedKoutra 2013: a similarity score for twographs on the same nodes, with attribution.
Lindsey et al., Sparse Crosscoders for Cross-Layer Features and Model Diffing, 2024. Koutra, Vogelstein, Faloutsos, DeltaCon, SDM 2013

Key results: what died, what was born, and same score, different delta

How to read: one bar = how many features sit at that position. Across: 0 = only in the model before the update, 0.5 = shared by both, 1 = only after. A feature is one concept the model uses (role play, refusal, step-by-step math).

The few changes say what the update did

Most features are shared. The few that die or are born say exactly what the update did, including what nobody asked for. A score cannot show this.

  • Crosscoder (Lindsey 2024): one dictionary of features read from both models at once
  • Born: refusal, correcting the assistant, step-by-step math. Died: role play, which nobody meant to remove
  • Same score, different delta (Shuttleworth 2025): LoRA and full fine-tuning tie on score, but LoRA adds new directions that drive forgetting
  • Duan 2026: probes break after updates, so every update needs rechecking
Lindsey et al., Sparse Crosscoders for Cross-Layer Features and Model Diffing, 2024. Shuttleworth et al., 2025. Duan et al., 2026

Critique of the tools, and the graph ancestor

The ancestor compared brains in 2013

Even the diffing tools can invent differences, so the comparison itself needs checking. Network science solved "compare two graphs on the same nodes" back in 2013.

  • Minder 2025: crosscoder artifacts can fake a diff signal
  • Kempf 2026: is white-box diffing worth its cost?
  • DeltaCon (this figure): two graphs on the same nodes, hubs weighted more, with attribution
  • Ours: a diff in the space of activation graphs, no dictionary to train

Remember: model diffing got its name in 2024, same score can hide a different delta, and DeltaCon already knew how to say which nodes moved.

Koutra, Vogelstein, Faloutsos, DeltaCon, SDM 2013, Figure 1

4. Eight surgeries, three lessons

Lesson 1Every change is one arrowa fine-tune, an edit, a steering pushfine-tune (task vector)knowledge editsteering directionone arrowthat changesthe modelMerging, editing and steering alladd the same kind of object.Ilharco 2023, Meng 2022, Arditi 2024Lesson 2Exam passed, inside bypassedthe old mechanism is still thereinoutolddetourexam: passedDPO, unlearning, editing, merging:the exam never asked the right question.Lee 2024, Hu 2025, Hammoud 2024Lesson 3Every inside score needs a nullwhat would chance score?observedchanceCKA, stitching, diffing: a score alonemeans nothing until compared with chance.Kornblith 2019, Smith 2025, Grรถger 2026So we need a tool that compares the inside before and after a change, against chance.Our paper builds the before/after part. The chance part is still open.
Ilharco et al., ICLR 2023. Meng et al., NeurIPS 2022. Arditi et al., NeurIPS 2024. Lee et al., ICML 2024. Hu et al., ICLR 2025. Hammoud et al., EMNLP Findings 2024. Kornblith et al., ICML 2019. Smith et al., ICML 2025. Grรถger et al., ICML 2026

5. The exam every surgery must pass is a sample

A benchmark is a set of questions. "Passed" means "passed what we thought to ask."

1. A benchmarkDrive trucks over the bridge.It holds: passed.That is the benchmark:it only sees whether the bridge held.2. A structural inspectionCheck which beams carry the load.Look at the structure, not only the result.For AI models we have onlylearned to drive trucks.3. Every family drives its own trucksEditingrewrite, paraphrase, localityMergingmulti-task accuracyPruningclassificationUnlearningforget versus retainWhy not look inside instead?Billions of unlabeled numbers.A full behavioral test costs morethan the change itself.And changes route around old mechanismsrather than delete them: the exam misses them.

5. Why we need a cheap screen

Every surgery is accepted by an exam that only looks at answers.

1. A checkup asks questionsAny pain?Sleep well?Eat well?It only hears the answers.2. A scan looks insideIt sees what the answers hide.3. What AI models havecheckupsvery goodscansalmost noneHence: a cheap screen.The pattern you have now seen eight times: exam passed, inside quietly broken3.1 Universalexaminside3.2 Stitchingexaminside3.3 Mergingexaminside3.4 Pruningexaminside3.5 Steeringexaminside3.6 Fine-tuneexaminside3.7 Editingexaminside3.8 DiffingexaminsideNext: four receipts, then what a scan could look like.

The whole field in one story

A large model is trained once, then changed over and over after release.

trained oncea large modelthen changed, over and over, after releasenew facta wrong fact iseditedexam: passedreads answers onlyinsidetwo models areaveragedexam: passedreads answers onlyinsidehalf of the layers aredeletedexam: passedreads answers onlyinsidedangerous knowledge isunlearnedexam: passedreads answers onlyinsideits state is nudgedwhile it runsexam: passedreads answers onlyinsideEvery changeis acceptedby an exam.The exam onlylooks at answers.This field keeps discovering the same thing: the exam passed, and the inside was quietly broken.

A quick bet

What fraction of "perfect" edits fail on a new wording?

1. An edit that looks perfectexample, illustrativeThe edit: "The Space Needle is in Seattle"changed to Paris"The Space Needle is in ..."exact promptParis"Where is the Space Needle located?"rephrased promptParis"The Golden Gate Bridge is in ..."unrelated fact unchangedSan Francisco (as before)All three pass. The benchmark calls it perfect.2. Your betA wording the checks never used:"Tourists at the Space Needle arestanding in which city?"How many perfect edits fail on such a wording?about 5%about 20%about 50%0%25%50%75%100%Three options on one line: where would you put it?56%the edited model answers:Seattlethe edit did not hold3. The answer56%each dot: 1% of the perfect edits987 of 1,761 edits with perfect scoresfail on the new wording.Two models, three editors.And editing is not special. Four receipts next, one per kind of change.
Benbrahim, Goghrod, Rashme, Ji, Fu, Zhang, Chen. Delta-Graph Signatures for Post-Edit Risk Screening in Knowledge Editing. SDM 2026

Four receipts: the exam said pass

Editing. 96.8% success when the model is fed the correct prefix word by word (teacher forcing), 38.5% when it writes freely. The test was measuring the test. Yang 2025

Unlearning. Harry Potter "forgotten". A little fine-tuning on harmless general text and the original comes back word for word. Hu 2025

Merging. A safe chemistry expert averaged with an unsafe math expert passes every capability benchmark and inherits the unsafe behavior. Hammoud 2024

Pruning. A quarter of the layers removed and repaired: classification keeps 80 to 90%, math and code generation drop hard and do not come back. Shrestha 2026

Same shape every time: the benchmark passed, a side effect hid, and only a different, costlier test found it. Can an internal signal find it first?

What is a delta graph?

Ask the same questions before and after an edit, and see how the model's parts now work together differently.

running example from slide 45: the Space Needle edit, Seattle to Paris1. The partslayer 1layer 2layer 3layer 4attentionMLPEach layer has an attentionblock and an MLP block (slide 2).These are the nodes:2 nodes per layer.2. Which parts work togetherWhere is the Space Needle?The Space Needle is located in ...The Golden Gate Bridge is in ...12 questions:8 about the fact,4 about unrelatedfactsblock Ablock Bquestion 1 to 12bar height = how active the block isrise and falltogether:draw an edgeTwo blocks that move together get an edge(a partial correlation, so shared drivers are removed).3. Before minus afterbeforevsafter=delta graphstrongerweakerSame 12 questions, same parts.Delta graph = after minus before:the edit's fingerprint inside the model.Why it helps: edits that will fail on a new wording leave a different fingerprint than edits that hold.Our paper learns which fingerprints are risky, and checks those edits first (next slides).
Benbrahim et al., SDM 2026

6. Our paper: Delta-Graph Signatures

1. Many editsAll three standard checks pass.The exam says: success.(1,761 edits in the paper)2. The truthAsk in a wording the checksnever used:56% fail.Finding them needs acostly extra test.3. Our screenNo extra test.Look only at how the insidechanged before vs after theedit (the delta graph).A risk model scores each edit.The risk model is trained once onedits whose outcome was tested(200 facts).4. Rank by riskcheck these firstHow we know it works: test the held-out wordings afterwards and see whether the top-ranked edits really fail.Top 20% by risk: 86% truly fail.The other 80%: 49%.A triage tool: it tells you what to check first, not a final verdict.Like a hospital triage desk: a few quick signs decide who sees the doctor first.Not a new editor and not a final verdict.Builds on DeltaCon (Koutra 2013): compare two graphs on the same nodes, before and after.
Benbrahim, Goghrod, Rashme, Ji, Fu, Zhang, Chen. Delta-Graph Signatures, SDM 2026. Koutra, Vogelstein, Faloutsos, DeltaCon, SDM 2013

Method: a connectome before and after the edit, in eight steps

MeditM'1Edit the modelM before, M' after12promptsper fact212 probe prompts per factdisjoint from edit andevaluation promptsattnMLP12 prompts2Lnodes3Record every blockmean activation of each attentionand MLP block: 2L nodes x 12 promptsfor Mand M'4Partial-correlation graphwhich blocks move togetherfMRI recipe, ridge, one fixed ฮปbefore the editafter the editafter minus beforeA real edit from the paper, steps 4 and 5. Each map is the graphas a grid: dark = two blocks move together. The edit shows as one spot.minus=ฮ”A5Subtractafter minus beforeฮ”A = A post minus A prehow muchconcentrateddepthspectral6A few featureshow much, how concentrated,where in depth, and spectral(graph Fourier, pre-edit graph)per (model, editor)7Rank-normalizecompare edits only withineach (model, editor)riskprior + delta graph8Score the riskL2 logistic regression ona model-and-editor prior
Benbrahim et al., SDM 2026, Figure 1.1 (panel b). Grouped CV by (model, fact). About 7.4 min per fact on one L40S

Results: same checking budget, more bad edits caught

100 edits, all with perfect scores. 56 are secretly bad. Budget: check only 20. Which 20?Pick at randomabout 11 bad foundPick by editor and model onlysome editors fail more oftenabout 14 bad foundPick by the delta graph (ours)how the inside changedabout 17 bad foundeach card = one edit checked. red cross = really bad, green tick = checked for nothing.Same 20 checks: 11, 14, 17. The extra catches come only from how the inside changed.(rates from the paper: 56%, 70%, 86% bad among the 20% checked first)0.703 to 0.797Main result: ranking quality (AUC)editor and model only, then with the delta graphAUC: 0.5 = chance, 1 = perfect ranking0.853 to 0.904Second dataset (wiki_recent)500 real Wikipedia facts, the same patternAUC: 0.5 = chance, 1 = perfect rankingNext: what carries the signal, the size of the change or its shape?
Benbrahim et al., SDM 2026, Tables 5.1 to 5.5

What carries the signal: size or shape?

Ablation: is it the size of the change, or its shape?Same probes, same before/after difference, scored different ways (AUC on the 1,761 benchmark-perfect edits)single numbers alone (no graph)0.52 to 0.59best 0.589: max edge changeeditor and model only (prior)0.705prior + best single number0.752+0.047prior + full delta graph (ours)0.798+0.0930.5 = chance0.60.70.8A single number barely beats chance. The graph's shape doubles the gain.So it is not how much the inside changed, but how the connections rearranged.Table F.1. The main text reports 0.703 and 0.797 for the same experiment. The five single numbers: max edge change 0.589, mean edge change 0.572,total change 0.575, top-5 node share 0.564, top-1 node share 0.521.Also holds: on 500 unseen facts, AUC 0.840 to 0.862. Rerun the 50 riskiest from scratch: only 2 of 50 flip, per model.
Benbrahim et al., SDM 2026, Table F.1, Tables 5.2 and 5.5

Limits, said first by the paper

Strong on Qwen, weaker on LlamaOn Llama, ROME and MEMIT edits thatpass the exam almost all fail, so there arehardly any good edits to tell apart. Anyscorer drops toward chance there (ROME0.449, MEMIT 0.498), even strongbehavioral ones.Still useful: the riskiest 20% on Llama are84% bad.0.847Qwen0.677Llama0.5AUCCheaper, not more accurate, than readingthe outputsSo: triage first, run the costly test only onthe flagged 20%, about 5 times lessfollow-up work.read the outputs: AUC 0.984costly per edit: free generationplus scoring on every editours: AUC 0.797reuses the probesCatches about a third of all failuresComputed from Table 5.1 (the paper doesnot state 30% directly).The value is in choosing what to checkfirst, not in catching everything.1,761 edits, 987 bad in totalcheck 20%: about 352about 300 of them badNothing to catch on CounterFactThose edits either fail the exam orsucceed deeply. The shallow failure wetarget does not occur there: a scopeboundary, not a counterexample.failure ratebelow 0.3%(1,110 edits, 200 facts)Also: the label depends on which 12 held-out prompts define failure. Scale: 7B and 8B models, three editors, two benchmark families.
Benbrahim et al., SDM 2026, Table 5.1, Figure 5.2, Sections 5.2 to 5.6

The closest competitors

what it readsfacts graphgradientsoutputshidden stateslearned featuresweightsactivation graphbefore the editaround the editafter the editwhen it looksBaser 2026, CLaREQin 2024, GradSimYang 2024, perplexityGu 2026Wang 2024, GIEYoussef 2025Kassem 2025, mnemeGupta 2025Zhang 2026ours: delta graphHow ours differsBaser predicts before, we screen after: two stages of onepipeline.GIE is a metric, ours is tested as a predictor of failure onheld-out wordings.GradSim needs gradients (a backward pass), ours does not.Youssef asks was it edited, we ask will it fail.mneme needs a trained dictionary, ours does not.Perplexity is the one-number ancestor (slide 51 ablation).
Baser 2026. Qin 2024. Yang 2024. Gu 2026. Wang 2024. Youssef 2025. Kassem 2025. Gupta 2025. Zhang 2026. Benbrahim et al., SDM 2026

Our paper, if you remember one line

56%of perfect edits failPass the three checks,still fail held-out prompts.987 of 1,761.The exam cannot see it.0.703 to 0.797AUC0.7030.797dashed: best single scalar, 0.5890.5prior+ delta graphTopology and spectrum of thebefore/after partial-correlationblock graph. No single scalarexceeds 0.589.top 20%screened firstcatches about 30% of hidden failuresCuts the expensive evaluationabout five-fold.Triage, not a verdict.validatedagainst behaviorfit on200 factstest on 500unseen factsfact-disjoint transfertrained on 200 facts, tested on 500unseen facts: AUC 0.840 to 0.862rerunprospective rerunsrerun the 50 riskiest edits from scratch:only 2 of 50 flip, per modelBoth show the signalis validated against behavior.
Benbrahim et al., SDM 2026

7. The map, revisited: what is filled and what is not

Green: a family-specific internal signal exists. Red: a cheap structural screen validated against held-out behavior exists for one family only.

Five gaps a network scientist can own

1One signature for all sevenfamiliesThe block-level delta graphis family-agnostic byconstruction. Untested formerging, pruning, unlearning.one delta graph2A calibrated null modelGraph two-sample tests andpermutation nulls exist.Nobody has applied them tobefore/after activationgraphs of a model.observedby chance3Attribution, not just detectionFrom "this edit is risky" to"these two attention blocksand this MLP moved."DeltaCon showed how in 2013.attnattnMLP4Suppression versus removalA bypass should add edgesaround a mechanism, adeletion should remove them.No statistic measures this yet.bypassedges addedremovaledges removed5Beyond textEditing and unlearning have reacheddiffusion models and are starting inmusic and speech. No structural screenexists for any of them. Ongoing work inmy group on locating and steeringmusical attributes in music generators isa natural test bed.diffusionmusicspeechThe three lessons in one breath: every change is one arrow, the exam can pass while the inside is bypassed,and every score read from inside the model needs a null model.

What to take home

For everyoneA benchmark is a sample."Passed" means "passed what we asked."Models are changed constantly after release,and every change deserves a look inside.orange: the questions the exam askedFor practitionersUse the before/after comparison.Every family now has a documented caseof hidden damage and an internal signalthat revealed it.beforeaftervswhatchangedinsideFor the graph peopleA trained model is a network.These are the tools this field is missing,and you already have them.null modelstwo-sample testschange attributionspectral summariesThank you.Questions and comments are welcome.
Esc or click a slide to return ยท click any slide to jump