FFA Theory

The Price of Locality:
Why Forward-Forward Underperforms Backpropagation?

A convergence and representation theory for understanding the performance gap between local learning and globally coordinated optimization.

Cornell Tech · Cornell University
NeurIPS 2026 · Poster

Geoffrey Hinton introduced the Forward-Forward Algorithm in 2022: it trains each layer through positive and negative forward passes with local objectives (original paper).

TL;DR

  • FFA underperforms BP. The performance gap persists across vision and language tasks and becomes more pronounced as task difficulty and model scale increase.
  • Failure mode 1: learning from shifting representations. Although one FFA block can be easy to optimize in isolation, earlier blocks keep changing the features received by later blocks during concurrent training.
  • Failure mode 2: representation collapse. Under the analyzed kernel-product conditions, representation kernels contract toward rank one with depth. For the collapsed-goodness-gradient subclass, local error signals then become increasingly redundant.
  • Design insight: preserve task-relevant directions. Local label readouts (LCE) give each block direct task supervision; spectral flattening spreads weight updates across more directions. Both narrow the measured gap to BP.

BP versus FFA

Backpropagation and Forward-Forward use the same forward computation but organize learning signals in fundamentally different ways. BP waits for the task loss, then transports one coordinated signal through the full chain rule. FFA attaches an independent contrastive objective to every layer and updates each block without gradients crossing layer boundaries.

Backpropagation

one task loss
one global backward path

Forward → store · backward ← release

Forward through L1; keep its activation for backward.1 / 6 stored

Forward-Forward

one objective per layer
no global backward path

Positive / negative passes → local supervision

L1 receives local supervision using one reusable buffer.1 buffer in use

Gray with a bar means activation is stored; green ✓ means updated. Hover or focus a layer to inspect it.

The animation illustrates when BP stores and releases activations and when FFA applies local updates. The buffer labels are schematic layer-buffer units, not measured GPU bytes; parameters, optimizer state, and temporary workspaces are excluded.

One cat image, two FFA passes

This supervised CIFAR-10 example uses two forward passes: cat photo + “cat” label is positive, and the same cat photo + “truck” label is negative. Here “negative” means an incorrect image–label pair, not a photograph of a truck. The photos and three-layer flow are illustrative, not actual CIFAR-10 samples or measured scores.

Positive pass · correct label
Orange cat photograph used as the input image CAT label

Same cat image + encoded “cat” label overlay → forward through the layers

L1: goodness ↑L2: goodness ↑L3: goodness ↑

This pass contributes a loss that pushes each layer's goodness above its threshold.

Negative pass · wrong label
The same orange cat photograph used in the positive pass TRUCK label

Same cat image + encoded “truck” label overlay → forward through the layers

Truck photo used only to illustrate the wrong class name Class cue only. A truck photo with the true label “truck” would be positive in its own example.
L1: goodness ↓L2: goodness ↓L3: goodness ↓

This pass contributes a loss that pushes each layer's goodness below its threshold.

Photos: cat by Halved sandwich and truck by Mattes (public domain).

Technical detail: formulations of BP and FFA

Forward inference

Let $x_i\in\mathbb R^{d^{(0)}}$ be sample $i$ and set its input representation to $h_i^{(0)}=x_i$. Suppressing the training-iteration index, the scaled residual network computes $$h_i^{(\ell)}=h_i^{(\ell-1)} +\frac{1}{L}\phi\!\left( \operatorname{LN}\!\left(W^{(\ell)}h_i^{(\ell-1)}\right) \right), \quad \ell=1,\ldots,L.$$ Thus $W^{(\ell)}\in\mathbb R^{d^{(\ell)}\times d^{(\ell-1)}}$ contains the trainable parameters of block $\ell$, while $h_i^{(\ell)}$ is the representation of the same sample after that block.

At training iteration $k$, we write these quantities as $W_k^{(\ell)}$ and $h_k^{(\ell)}(x_i)$. Stacking the representations of all $N$ samples row by row gives $$H_k^{(\ell)} =\begin{bmatrix} h_k^{(\ell)}(x_1)^\top\\[-2pt] \vdots\\[-2pt] h_k^{(\ell)}(x_N)^\top \end{bmatrix}\in\mathbb R^{N\times d^{(\ell)}}$$ and $W_k=(W_k^{(1)},\ldots,W_k^{(L)})$ denotes the full network.

Backpropagation

Every layer differentiates the same end-to-end task loss:

$$g_{k,\mathrm{BP}}^{(\ell)} =\nabla_{W^{(\ell)}}\mathcal L_{\mathrm{task}}(W_k), \qquad W_{k+1}^{(\ell)} =W_k^{(\ell)}-\eta g_{k,\mathrm{BP}}^{(\ell)}.$$

The chain rule lets the final task error coordinate the updates of all layers.

Forward-Forward

Block $\ell$ differentiates only its own local loss and treats the incoming representation $H_k^{(\ell-1)}$ as fixed:

$$g_{k,\mathrm{FFA}}^{(\ell)} =\nabla_{W^{(\ell)}}\mathcal L^{(\ell)} \bigl(W_k^{(\ell)};\operatorname{stopgrad}(H_k^{(\ell-1)})\bigr), \qquad W_{k+1}^{(\ell)} =W_k^{(\ell)}-\eta^{(\ell)}g_{k,\mathrm{FFA}}^{(\ell)}.$$

No task gradient crosses a block boundary, so these updates are not gradient steps on one shared global objective.

FFA underperforms BP

The starting point of this work is empirical: FFA remains substantially less accurate than BP even when both are trained on the same fixed backbone and dataset. In the strictly local endpoint of our CNN12/CIFAR-10 intervention, the gap is 32.65 percentage points.

Central question. What mechanisms cause FFA to underperform backpropagation?

We answer this question by identifying two distinct but coupled failure modes: concurrent optimization on shifting representations and geometric representation collapse. The later experiments test how the gap changes with locality, model scale, and update geometry.

Two coupled failure modes

1. Learning from a moving representation

Each FFA block receives its inputs from the blocks before it. Because all blocks train at the same time, those inputs change whenever an earlier block updates. A downstream block can reduce its local loss for the current representation, only to receive a different representation at the next step. Thus, every block may make local progress without the network following one coordinated descent direction.

Why the downstream target moves
earlier layers update $W_k^{(<\ell)}\to W_{k+1}^{(<\ell)}$ upstream network changes
layer $\ell$ sees new features $H_k^{(\ell-1)}\to H_{k+1}^{(\ell-1)}$ its input has moved
its local loss changes $\mathcal L_k^{(\ell)}\to\mathcal L_{k+1}^{(\ell)}$ the target has moved

Because $\mathcal L_k^{(\ell)}(W)=\mathcal L^{(\ell)}(W;H_k^{(\ell-1)})$, layer $\ell$ is not repeatedly descending one fixed loss: every upstream update changes the loss it sees.

What this does to one local update

A local objective whose optimum moves between iterations The downstream update moves toward the old local optimum, but an upstream representation change shifts the loss contours and their optimum to a new location. Wₖ at layer ℓ update using the old loss old contours new contours old optimum new optimum

The gradient can be a valid descent direction for $\mathcal L_k^{(\ell)}$ without being aligned with the shifted objective $\mathcal L_{k+1}^{(\ell)}$.

FFA's local step and the upstream feature update are each reasonable in isolation; their simultaneous coupling makes the downstream objective move between steps.

The multi-layer theorem captures this interference through a linearly decaying optimization term plus a coupling residual.

Key theorem: multi-layer FFA convergence with coupling residual

Inside the paper's local regularity regime, each isolated block satisfies a PL inequality. For concurrent training, define the average local suboptimality $$V_k=\frac{1}{L}\sum_{\ell=1}^{L} \left(\mathcal L_k^{(\ell)}-\mathcal L^{(\ell)*}\right).$$ Under the boundedness, smoothness, Gram-conditioning, sensitivity, scaled-residual, and step-size conditions stated in the paper, $$V_k\le \left(1-\frac{\eta\mu_{\min}}{2}\right)^kV_0 +O\!\left(\frac{e^{2\rho}\beta_{\max}}{\mu_{\min}^2}\right).$$ The first term is the progress inherited from the local PL geometry. The second controls interference caused by simultaneous changes in upstream representations.

2. Geometric representation collapse

Representations become alike

One intuition is that each forward block can slightly reduce differences between samples' feature directions. Repeating such blocks can accumulate these small reductions, making initially distinct representations increasingly alike. Consequently, the local goodness gradients become increasingly redundant, and the network loses its ability to distinguish between classes.

Imagine comparing every pair of training images after each layer. Let $h_i^{(\ell)}$ be the representation (a.k.a. activation or feature) of sample $i$ after layer $\ell$, and let $\widetilde h_i^{(\ell)}=\operatorname{LN}(h_i^{(\ell)})$ be its normalized version. The representation kernel $\Sigma^{(\ell)}$ is the resulting $N\times N$ similarity table: entry $(i,j)$ compares images $i$ and $j$ at depth $\ell$. Using the representation matrix $H^{(\ell)}$ defined above (with the iteration index suppressed), Then $$[\Sigma^{(\ell)}]_{ij} =\frac{\langle\widetilde h_i^{(\ell)},\widetilde h_j^{(\ell)}\rangle}{d}, \qquad \Sigma^{(\ell)}\in\mathbb R^{N\times N}.$$ Intuitively, $\Sigma^{(\ell)}$ can evaluate how similar the network's representations are for different samples. In the extreme case, if all normalized representations point in the same direction, $\Sigma^{(\ell)}$ becomes the all-ones matrix $J_N$. The paper shows that this stronger contraction can happens when the specified kernel-product conditions hold. At that time, it becomes hard to distinguish between samples.

Representation kernel contraction across depth A diverse similarity matrix progressively becomes an approximately uniform rank-one matrix as network depth increases. shallow diverse similarities depth deeper directions merge rank one Σ(ℓ) → Jₙ
Under the paper's kernel-product conditions, similarities can converge toward the rank-one matrix $J_N$ with depth, including similarities between different classes.
Key theorem: convergence toward the rank-one kernel

Under the kernel-product conditions (see the formal definition in the paper), if the leading Lyapunov exponent satisfies $\lambda_1<0$, then $$\left\|\Sigma^{(\ell)}-J_N\right\|_F \le \mathcal O\!\left(e^{-|\lambda_1|\ell/2}\right) \left\|\Sigma^{(0)}-J_N\right\|_F.$$ The representation kernel therefore approaches $J_N$ geometrically with depth.

Why BP can preserve distinct learning signals

Imagine that cat, dog, and truck images have almost the same features at one layer. The network gives them similar predictions, although their correct labels call for different corrections.

FFA: Each layer forms a new signal from its local goodness objective. If those activations align, their corrections also point along nearly the same line; positive and negative examples may reverse the sign, but add no new direction. This is the analyzed collapsed-gradient case, not a property of every local goodness.

BP: The final loss uses each image's label. With three classes, the resulting output errors can span more than one direction even when the predictions are nearly identical. Backward Jacobians carry these errors to earlier layers; if they preserve those directions, BP can send varied corrections despite similar current features. A Jacobian can also suppress directions, so this is a possibility rather than a guarantee.

$\Sigma$ measures similarity among the features. To see whether the corrections have also become redundant, we use their second-moment matrix $\Gamma$: effective rank near one means that most correction energy lies along one line. Effective learning capacity (ELC) sums the effective rank above this one-direction baseline across layers. It measures the diversity of available learning signals, not a guarantee of task accuracy. In this paper, we show that ELC of BP can increase with depth, while ELC of FFA can remain near one, which also explains the accuracy gap.

Technical detail: $\Gamma$, effective rank, and ELC

Let $\delta^{(\ell),\mathcal A}(x)$ be the error signal algorithm $\mathcal A$ sends to layer $\ell$. Its second-moment matrix is $$\Gamma^{(\ell),\mathcal A} =\mathbb E_x\!\left[\delta^{(\ell),\mathcal A}(x)\delta^{(\ell),\mathcal A}(x)^\top\right].$$ For a positive semidefinite matrix $M$, define $$\operatorname{erank}(M)=\frac{(\operatorname{tr}M)^2}{\operatorname{tr}(M^2)},\qquad \operatorname{ELC}(\mathcal A)=\sum_{\ell=1}^{L}\left(\operatorname{erank}(\Gamma^{(\ell),\mathcal A})-1\right).$$ Subtracting one removes the rank-one baseline at each layer. The measure does not assert that directions from different layers are mutually orthogonal.

Key theorem: effective learning capacity

Under the corresponding BP and FFA conditions established in the paper,

$$\operatorname{ELC}(\mathrm{BP})=\Omega(L), \qquad \operatorname{ELC}(\mathrm{FFA})=O(1).$$

Takeaway. When representation collapse limits FFA, narrowing its gap to BP calls for preserving task-relevant differences between samples as depth grows. Two possible interventions give local learning more ways to retain those differences.

Potential repairs: richer local targets and flatter updates

Give each block a supervised readout

One route gives every block more label information: attach a trainable readout to its output and train that block with cross-entropy against the true label or next token. This is Local Cross-Entropy (LCE). Like FFA, LCE stops gradients at block boundaries, but each block now receives a task-label signal through its own readout rather than a positive/negative goodness signal. In the language-model scale sweep, each block has its own vocabulary readout; the displayed parameter counts cover the shared Transformer backbone and exclude these extra heads.

Spread each weight update across more directions

Another route changes the update geometry directly. For an update $\Delta W=U\operatorname{diag}(s_i)V^\top$, interpolate its squared singular values toward their mean: $\lambda_i(p)^2=(1-p)s_i^2+p\overline{s^2}$. This keeps the update's Frobenius norm fixed while flattening its spectrum as $p$ increases. The mechanism test below applies this rule to convolutional updates, then measures whether the error-signal spectrum $\Gamma$ and task accuracy improve.

Explore the measured evidence

The controls below expose the measured sweeps behind the headline results. In the network diagrams, orange dots mark where each block receives local supervision, while the blue dot marks the task loss at the final layer. Orange segments show separate local gradients; the blue reference shows the end-to-end BP pathway.

Vision: controlled locality intervention

The task is ten-class CIFAR-10 image classification with a fixed 12-layer convolutional network. Grouped FFA trains on true-label positive and wrong-label negative overlays; the slider changes how many contiguous layers share each local goodness objective. Architecture and training protocol stay fixed, so this sweep isolates the block partition.

CNN12 on CIFAR-10

The backbone always has 12 layers. Only the number of independently trained contiguous blocks changes.

Best accuracy53.02%
Final local loss1.1014
Mean erank Γ2.198

What the diagram shows. Fewer block boundaries let each goodness signal assign credit across more of the same 12-layer backbone. From 12 local blocks to one global block, mean accuracy rises from 53.02% to 78.74%, although BP remains higher at 85.67%. The intervention links stricter locality to weaker learning on this fixed task.

Blocks × layersBest accuracyFinal lossMean erank(Γ)
12 × 153.02 ± 0.48%1.1014 ± 0.00342.198 ± 0.042
6 × 258.09 ± 1.85%1.0455 ± 0.00685.146 ± 0.592
4 × 367.57 ± 4.35%1.0064 ± 0.01584.856 ± 0.552
3 × 472.22 ± 1.07%0.9795 ± 0.00185.135 ± 2.164
2 × 676.21 ± 0.48%0.9657 ± 0.00277.613 ± 3.828
1 × 1278.74 ± 4.06%0.9550 ± 0.01997.975 ± 1.214
BP85.67 ± 0.09%5.87e−5 ± 7.13e−531.654 ± 3.793

Measured means ± sample standard deviation over three seeds. The selected row is highlighted.

Language modeling: a scale sweep

The task is next-token prediction on an OpenWebText subset with 50 million training tokens and 0.5 million validation tokens. Decoder-only Transformers range from 2-layer Tiny to 12-layer XLarge, use GPT-2 tokenization and 256-token contexts, and are compared by validation perplexity under scale-dependent token budgets. The slider selects a model size; it does not change how a fixed model is partitioned.

Transformer pre-training on OpenWebText

For vanilla FFA, every Transformer layer is its own local block. Model depth, parameter count, and compute budget all increase across this sweep.

BP perplexity28
FFA / BP9.19×
LCE / BP1.51×

What the diagram shows. Vanilla FFA keeps one local objective per layer as the model grows; LCE uses the same gradient-isolated layout with cross-entropy readouts. Across the measured scales, FFA/BP perplexity grows from 3.73× to 9.19× and LCE/BP from 1.20× to 1.51×. LCE narrows the gap, but neither local method matches BP; this scale sweep does not isolate depth as the cause.

ScaleLayersParametersBP PPLFFA PPLLCE PPLFFA/BPLCE/BP
Tiny27M1385161663.73×1.20×
Small416M68313844.59×1.22×
Medium630M48275645.74×1.33×
Large851M39256546.55×1.38×
XLarge12124M28258429.19×1.51×

This slider selects measured model scales. Because depth, width, parameter count, and compute change together, the trend is not a fixed-architecture causal ablation of locality.

Mechanism test: flattening the update spectrum

On CNN3/CIFAR-10, we apply the spectral interpolation described above to each convolutional update. The backbone, data, each method's objective and optimizer, and the training budget stay fixed across $p$.

CNN3 on CIFAR-10

$p=0$ is the original update; $p=1$ replaces its singular values by a flat spectrum with the same Frobenius norm.

FFA accuracy57.74%
FFA erank(Γ)4.42
Gap to BP21.23 pp
Update singular-value spectrum · illustrativeUpdate PR = 8.00 / 8

Bars show singular values of one eight-direction example with fixed Frobenius norm. Update PR is the participation ratio of their squared values.

FFA error-signal rank · measurederank(Γ) = 4.42

Points are terminal mean layerwise erank(Γ), averaged over three seeds. The highlighted point follows the slider.

The left spectrum illustrates the exact interpolation rule on an example; its bars are not measured CNN update singular values. The right plot uses the measured error-signal ranks in the table. These ranks describe different matrices.

What the plots show. Increasing $p$ spreads update energy across more directions by construction. The measured FFA error-signal rank rises overall, though not monotonically, and accuracy improves by 10.34 points from $p=0$ to $p=1$. A 21.23-point gap to BP remains at $p=1$, so flattening helps without closing the locality gap.

Flattening $p$BP accuracyFFA accuracyBP erank(Γ)FFA erank(Γ)
0.077.56 ± 0.75%47.40 ± 0.98%42.78 ± 2.273.43 ± 0.49
0.278.80 ± 0.39%53.41 ± 1.25%50.75 ± 2.043.85 ± 0.24
0.479.08 ± 0.17%55.30 ± 1.21%53.01 ± 0.754.27 ± 0.43
0.679.11 ± 0.27%55.87 ± 1.65%53.85 ± 3.613.95 ± 0.83
0.879.13 ± 0.43%57.27 ± 1.15%50.84 ± 1.384.49 ± 0.64
1.078.97 ± 0.29%57.74 ± 0.50%44.48 ± 1.964.42 ± 0.61

Flattening improves FFA by 10.34 points, providing intervention evidence that update geometry matters. The remaining 21.23-point gap shows that geometry alone does not remove the cost of locality.

What should better local learning preserve?

The diagnosis suggests that improving the scalar goodness alone is not enough. A scalable local method must optimize locally without losing the distinctions needed by the global task.

  • Stabilize the target seen by each block. Limit upstream drift or explicitly track it during concurrent updates.
  • Allow selective cross-block communication. Even a limited global correction can coordinate otherwise independent objectives.
  • Preserve diverse learning directions. Monitor representation and error-signal effective rank, not only local loss.
  • Evaluate the actual task. Local convergence is an optimization diagnostic, not a substitute for downstream accuracy or perplexity.

BibTeX

@inproceedings{wu2026price,
  title     = {The Price of Locality: Why Forward-Forward Underperforms Backpropagation?},
  author    = {Wu, Zhaoxian and Liu, Haichuan and Chen, Tianyi},
  booktitle = {Advances in Neural Information Processing Systems},
  year      = {2026},
  url       = {https://arxiv.org/abs/2609.33240}
}