Geometry Conflict: Explaining and Controlling Forgetting in LLM Continual Post-Training

ABSTRACT

Continual post-training aims to extend large language models (LLMs) with new knowledge, skills, and behaviors, yet it remains unclear when sequential updates enable capability transfer and when they cause catastrophic forgetting. Existing methods mitigate forgetting through sequential fine-tuning, replay, regularization, or model merging, but offer limited criteria for determining when incorporating new updates is beneficial or harmful. In this work, we study LLM continual post-training through three questions: *What drives forgetting? When do sequentially acquired capabilities transfer or interfere? How can compatibility be used to control update integration?* We address these questions through *task geometry*: we represent each post-training task by its parameter update and study the covariance geometry induced by the update. Our central finding is that: **forgetting can be considered as a state-relative update-integration failure, it arises when the covariance geometries induced by tasks misalign with the geometry of the evolving model state**. Sequential updates transfer when they remain compatible with the model state shaped by previous updates, and interfere when state-relative geometry conflict becomes high. Motivated by this finding, we propose **G**eometry-**C**onflict **W**asserstein **M**erging (GCWM), a data-free update-integration method that constructs a shared Wasserstein metric via Gaussian Wasserstein barycenters and uses geometry conflict to gate geometry-aware correction. Across Qwen3 0.6B--14B on domain-continual and capability-continual settings, GCWM consistently outperforms data-free baselines, improving retention and final performance without replay data. These results identify geometry conflict as both an explanatory signal for forgetting and a practical control signal for LLM continual post-training.

Continual Post-TrainingModel MergingGeometry Conflict
2026年5月13日
Yuanyi Wang, Yifan Yang, Su Lu, Yanggan Gu, Pengkai Wang, Wenjun Wang, Zhaoyi Yan, Congkai Xie, Jianmin Wu, Jialun Cao, Shing-Chi Cheung, Hongxia Yang

Abstract

Continual post-training aims to extend large language models (LLMs) with new knowledge, skills, and behaviors, yet it remains unclear when sequential updates enable capability transfer and when they cause catastrophic forgetting. Existing methods mitigate forgetting through sequential fine-tuning, replay, regularization, or model merging, but offer limited criteria for determining when incorporating new updates is beneficial or harmful. In this work, we study LLM continual post-training through three questions: What drives forgetting? When do sequentially acquired capabilities transfer or interfere? How can compatibility be used to control update integration? We address these questions through task geometry: we represent each post-training task by its parameter update and study the covariance geometry induced by the update. Our central finding is that: forgetting can be considered as a state-relative update-integration failure, it arises when the covariance geometries induced by tasks misalign with the geometry of the evolving model state. Sequential updates transfer when they remain compatible with the model state shaped by previous updates, and interfere when state-relative geometry conflict becomes high. Motivated by this finding, we propose Geometry-Conflict Wasserstein Merging (GCWM), a data-free update-integration method that constructs a shared Wasserstein metric via Gaussian Wasserstein barycenters and uses geometry conflict to gate geometry-aware correction. Across Qwen3 0.6B--14B on domain-continual and capability-continual settings, GCWM consistently outperforms data-free baselines, improving retention and final performance without replay data. These results identify geometry conflict as both an explanatory signal for forgetting and a practical control signal for LLM continual post-training.

1. Introduction

Continual post-training is becoming an increasingly important paradigm for extending large language models (LLMs). Rather than learning jointly over all desired capabilities or data, a model is expected to learn through a sequence of post-training stages, each targeting a new domain, skill, or behavior. This process is natural for real scenarios as capabilities are introduced incrementally. However, sequential post-training faces a fundamental challenge: learning a new task undermines the knowledge acquired from previous ones, a phenomenon known as catastrophic forgetting, often driven by interference between sequential parameter updates.

Existing approaches can be broadly categorized into four classes: sequential fine-tuning, replay-based methods that revisit past data, regularization methods that constrain update drift, or model merging strategies that combine task-specific adaptations. These approaches have led to important progress, but they still lack a principled account of task compatibility in continual post-training. As a result, they often struggle to answer a central practical question: when should new parameter updates be strongly integrated into the current model, and when should such integration be restrained? This issue is particularly pronounced for LLMs, where tasks are highly heterogeneous, post-training objectives differ substantially, and the same update magnitude can lead to very different retention outcomes.

To address this problem, we study LLM continual post-training through three questions: What drives forgetting? When do sequentially acquired capabilities transfer or interfere? How can compatibility be used to control update integration? We answer these questions through a task-geometry view of post-training updates. Specifically, we represent each task by its parameter update and study the induced covariance geometry, which captures not only update magnitude but also the subspaces and spectral structure through which a task changes the model. We define geometry conflict as a normalized Bures--Wasserstein discrepancy between task-induced covariance geometries in a shared space, and use its state-relative form to measure compatibility with the evolving LLM state.

Our analysis (Sec. 3) across Qwen3 scales and continual strategies compares geometry conflict with update norm, subspace alignment ratio, and gradient conflict. It reveals a central mechanism: forgetting can be considered as a state-relative update-integration failure, it arises when the covariance geometries induced by tasks misalign with the geometry of the evolving model state, whereas transfer occurs when new updates remain compatible with the state shaped by previous updates. This explains why raw update norm and isolated pairwise compatibility are insufficient, and why geometry conflict serves as a natural signal for controlling sequential update integration.

Motivated by this finding, we propose Geometry-Conflict Wasserstein Merging (GCWM), a data-free update-integration method for LLM continual post-training. GCWM constructs task-induced covariance geometry, builds a shared Wasserstein metric via Gaussian Wasserstein barycenters, and uses geometry conflict to gate geometry-aware correction, which allows GCWM to perform compatibility-controlled update integration. We further provide theoretical support showing that the induced loss change is controlled by geometry conflict and gated merge displacement.

Across domain-continual and capability-continual settings, GCWM consistently improves retention and final performance over data-free baselines without replay data. On Qwen3 models from 0.6B to 14B, GCWM remains the strongest data-free update-integration method across scales, showing that geometry conflict is useful not only as an explanatory signal for forgetting but also as a practical control signal for continual post-training. In summary, our contributions are summarized as follows:

\ (i) We develop a task-geometry analysis of LLM continual post-training and show that forgetting is better explained as a state-relative update-integration failure, beyond update norm and isolated pairwise compatibility. \ (ii) We introduce geometry conflict, a Bures--Wasserstein distance over task-induced covariance geometries, and identify it as both an explanatory signal for forgetting and a compatibility signal for update integration, complementing existing subspace alignment ratio and gradient conflict. \ (iii) We propose Geometry-Conflict Wasserstein Merging, a data-free update-integration method that constructs a shared Wasserstein metric and gates geometry-aware correction by layer-wise conflict. \ (iv) We derive a conflict-controlled theory linking GCWM's relative loss to geometry conflict and gated merge displacement, and validate GCWM on Qwen3 0.6B--14B across domain- and capability-continual settings, improving final performance over data-free baselines without replay data.

2. Preliminary

2.1 Problem Setup

We study continual post-training for LLMs. Starting from a pretrained model with parameters θpre\theta_{\mathrm{pre}}, the model is adapted through a sequence of tasks T={T1,…,TK}\mathcal{T}=\{T_1,\ldots,T_K\}, where each task introduces a new domain, skill, or behavior. For task TtT_t, we denote its task-specific update by

Δt=θt−θpre,\Delta_t = \theta_t - \theta_{\mathrm{pre}},

where θt\theta_t is the model adapting to TtT_t. We use these task updates, which may be parameter-efficient or full-model updates, as the basic objects for analyzing in LLM continual post-training.

2.2 Task Geometry and Compatibility Signals

A task update is not fully characterized by its norm: two updates with similar magnitude can affect different subspaces and induce different forgetting behavior. For a layer ℓ\ell, let Δt(ℓ)∈Rdout×din\Delta_t^{(\ell)} \in \mathbb{R}^{d_{\mathrm{out}}\times d_{\mathrm{in}}} denote the update matrix of task TtT_t. Motivated by the task vector, we define task geometry as:

Ct(ℓ)=(Δt(ℓ))⊤Δt(ℓ),C_t^{(\ell)} = \big(\Delta_t^{(\ell)}\big)^\top \Delta_t^{(\ell)} ,

which captures the dominant directions of the update. Let Vi(ℓ),Si(ℓ)V_i^{(\ell)},S_i^{(\ell)} be retained SVD factors and Q(ℓ)=orth([Vi(ℓ),Vj(ℓ)])Q^{(\ell)}=\mathrm{orth}([V_i^{(\ell)},V_j^{(\ell)}]) their shared union basis. The projected geometry Bi(ℓ)=(Q(ℓ))⊤Vi(ℓ)(Si(ℓ))2(Vi(ℓ))⊤Q(ℓ)+λIB_i^{(\ell)}=(Q^{(\ell)})^\top V_i^{(\ell)}(S_i^{(\ell)})^2(V_i^{(\ell)})^\top Q^{(\ell)}+\lambda I is the regularized projection of the retained component of Ci(ℓ)C_i^{(\ell)}, with λ>0\lambda>0; Bj(ℓ)B_j^{(\ell)} is constructed analogously. Here, dBd_{\mathrm B} denotes the Bures--Wasserstein distance between positive-semidefinite matrices, and tr(A)=∑kAkk\mathrm{tr}(A)=\sum_k A_{kk} is the matrix trace. To compare two tasks, we project them into a shared basis and measure their discrepancy using a normalized Bures--Wasserstein distance :

γij(ℓ)=dB2(Bi(ℓ),Bj(ℓ))tr(Bi(ℓ))+tr(Bj(ℓ))+ε,\gamma_{ij}^{(\ell)} = \frac{ d_{\mathrm{B}}^2(B_i^{(\ell)},B_j^{(\ell)}) }{ \mathrm{tr}(B_i^{(\ell)})+\mathrm{tr}(B_j^{(\ell)})+\varepsilon },

where Bi(ℓ)B_i^{(\ell)} and Bj(ℓ)B_j^{(\ell)} are the projected geometries. We refer to γij(ℓ)\gamma_{ij}^{(\ell)} as geometry conflict. Lower values indicate more compatible task-induced geometries. In Sec. 3, we compare geometry conflict with three standard diagnostics: update norm, subspace alignment ratio (SAR), and gradient cosine conflict. State-relative variants replace one task update with the current continual-training state. Full metric definitions and aggregation details are provided in Appendix E.

Continual Post-training has become an increasingly important paradigm for extending LLMs beyond their original pretraining distribution, including domain adaptation, capability acquisition, and behavior alignment over sequential stages. Existing approaches largely follow four lines. Sequential fine-tuning directly adapts the model stage by stage, but is highly prone to forgetting under heterogeneous task sequences. Replay-based methods mitigate forgetting by revisiting historical data, while regularization-based methods constrain update drift to preserve prior knowledge. Model merging that combines task-specific adaptations offers a plug-in workflow, but struggles to resolve cross-task interference. However, most existing methods emphasize preserving prior performance during sequential updates while offering limited guidance on the task-compatibility conditions under which sequential interactions should be encouraged or suppressed. Our work addresses this gap through a task-compatibility perspective.

Continual Model Merging provides a data-efficient alternative to standard sequential adaptation by composing task-specific parameter updates in weight space. Recent work studies sequential settings in which models arrive incrementally over time, including projection-based sequential merging, stability-based methods based on null-space filtering or test-time gating, resource-constrained online merging of adapters, and broader hybrid frameworks that combine continual learning and model merging. Our method is instantiated as a data-free continual merging method, but the broader goal is to study continual post-training through task compatibility and use merging as an mechanism for exploiting the resulting compatibility findings.

Compatibility Metrics and Signals. Recent work studies compatibility via parameter discrepancy, gradient alignment, and subspace or spectral overlap. Demystifying Mergeability shows that subspace overlap and gradient alignment are stable method-agnostic indicators, but these signals remain largely diagnostic. In contrast, we introduce geometry conflict as a method-native control signal derived from task-induced covariance geometry, and construct a shared merging metric via Bures--Wasserstein geometry and Gaussian Wasserstein barycenters.

3. What Governs Forgetting in Continual Post-Training?

Before introducing GCWM, we first ask what makes a continual post-training step harmful. Across Qwen3 models from 0.6B to 14B and four representative strategies--Seq.\ SFT, EWC regularization, FOREVER replay, and AIMMerging --we compare forgetting with update norm, SAR, gradient conflict, and our geometry conflict. Here, retention loss is the positive old-task drop from each task's best previous score, reported in percentage points (pp) when scaled by 100, and ρs\rho_s denotes Spearman rank correlation. The analysis yields four findings: update norm is only a coarse drift baseline; geometry conflict refines SAR-based compatibility; state-relative geometry mismatch best tracks continual forgetting; geometry and gradient conflict reveal complementary failure modes. Extended diagnostics and bootstrap confidence intervals are organized in Appendix F.

3.1 Update Norm Is Insufficient to Explain Forgetting

Figure 1. State-relative geometry tracks forgetting across continual steps and scales. Panel (a) shows normalized SFT dynamics within each scale; panel (b) reports ∣ρs∣\vert \rho_s\vert between each signal and loss.

A natural hypothesis is that forgetting is mainly driven by parameter drift: larger updates should induce larger retention loss. We test this by comparing update norm with forgetting, and contrast it with geometry signals that use different reference points. In Figs. 1 and 2, active conflict is the mean pairwise geometry conflict among active task updates, while state and global gaps measure geometry mismatch between active task updates and the evolving model state. Fig. 2(a) shows that update norm has a nontrivial but coarse association with retention loss (∣ρs∣=0.48\vert \rho_s\vert =0.48). State-relative geometry is stronger: the global state-active gap reaches ∣ρs∣=0.59\vert \rho_s\vert =0.59, exceeding both update norm and active-pair conflict (∣ρs∣=0.30\vert \rho_s\vert =0.30).

Figure 2. Global and method-level associations. Top: global ∣ρs∣\vert \rho_s\vert . Bottom: signed method-level ρs\rho_s; FVR denotes FOREVER.

The scale breakdown in Fig. 1(b) further shows that this advantage becomes clearer in larger LLMs: the global gap increases from 0.160.16 at 0.6B to 0.860.86 at 14B, while update norm remains a weaker drift baseline. Overall, update norm measures how far the model moves, but not whether the movement remains compatible with task-induced geometries. Bootstrap confidence intervals and additional step-level rankings are provided in Appendices F.1 and F.3.

3.2 Geometry Conflict Refines Subspace Compatibility

Subspace overlap is a natural compatibility proxy: if two updates act on similar directions, they may be easier to integrate. We therefore compare SAR with geometry conflict (Sec. 2.2). As shown in Fig. 3(a), SAR and geometry conflict are related but non-redundant: their global rank association is moderate (ρs=0.27\rho_s=0.27), and task pairs with similar SAR can still exhibit very different geometry conflict. SAR captures where updates overlap; geometry conflict captures whether their induced covariance geometry is compatible in that shared space.

Pairwise geometry is useful for regime diagnosis, but it is not a standalone predictor of forgetting. In Fig. 3(a--c), SAR percentile ranks task-pair SAR values; GC-drop and GC-forget denote correlations between pairwise geometry conflict and immediate old-task score change or best-previous forgetting, respectively. Fig. 3(b) shows that GC-drop stays near zero across methods and scales, while GC-forget is scale-sensitive: it is visible on 0.6B--4B (0.28/0.31/0.300.28/0.31/0.30) but weak on 8B and 14B (0.12/0.020.12/0.02). Fig. 3(c) further illustrates this point: large drops, such as Math→\rightarrowHistory (12.912.9 pp) and Math→\rightarrowEconomics (12.312.3 pp), do not form a single pairwise-conflict pattern. Thus, pairwise compatibility is informative but insufficient, motivating the state-relative analysis in Sec. 3.3. Overall, SAR and geometry conflict capture different levels of compatibility. Pairwise confidence intervals, heatmaps, summaries, and harmful transitions are provided in Appendices F.1 and F.4.

Figure 3. Pairwise compatibility is informative but insufficient. (a)--(c) SAR and geometry conflict stratify task-pair transfer regimes, while pairwise conflict alone weakly predicts forgetting. GC-drop is the signed association with the immediate old-task delta; GC-forget measures degradation from each old task's best prior score.

3.3 State-Relative Geometry Conflict Tracks Continual Forgetting

Sec. 3.2 shows that isolated task pairwise compatibility is incomplete. In LLM continual post-training, each incoming update is applied to an evolving model state that already encodes previous updates. The question is whether incoming task geometry remains compatible with the current state.

Fig. 1(a) tracks this effect under Seq.\ SFT. Active-pair conflict fluctuates across steps, while state and global gaps more closely follow the growth of retention loss, especially from 1.7B to 14B. The method-level heatmap in Fig. 2(b) shows the same pattern is strongest under direct sequential updating: state/global signals reach 0.68/0.700.68/0.70 for Seq.\ SFT and remain substantial for EWC (0.40/0.420.40/0.42), but weaken when replay or merging compresses forgetting variance. This identifies the evolving model state, rather than isolated task pairs, as the relevant reference point for geometry-based forgetting analysis. Full confidence intervals and method-stratified correlations are in Appendices F.2--F.3.

3.4 Geometry and Gradient Conflict Reveal Complementary Failure Modes

Finally, we ask whether geometry conflict simply duplicates gradient conflict. The answer is no. Here, q/k/v/oq/k/v/o denote attention projections, gate/up/downgate/up/down denote MLP projections, top-layer share is the fraction of top-ranked conflict layers in each family, min grad-cos is the minimum gradient cosine, and neg-grad ratio is the fraction of negative-cosine pairs. Fig. 4(a) shows a sharp module-level separation: top geometry-conflict layers concentrate in up_projup\_\mathrm{proj}, gate_projgate\_\mathrm{proj}, v_projv\_\mathrm{proj}, and down_projdown\_\mathrm{proj}, whereas top negative-gradient layers are dominated by k_projk\_\mathrm{proj} and q_projq\_\mathrm{proj}. Together, the four geometry-heavy families account for about 0.890.89 of top geometry-conflict layers, while query/key projections account for about 0.860.86 of top negative-gradient layers. Fig. 4(b) further shows that the geometry-conflict locus changes with the update-integration strategy, while negative-gradient conflict remains consistently query/key-centric. Fig. 4(c) complements this module view: the global geometry gap is the strongest forgetting-aligned signal among the plotted predictors, whereas gradient diagnostics are more aligned with old-task mean and overall performance. Thus, geometry conflict and gradient conflict are complementary diagnostics. Gradient conflict exposes optimization-level opposition, while geometry conflict captures update-integration mismatch. This distinction is important for GCWM: geometry conflict is not used as a replacement for gradient diagnostics, but as a native signal for controlling how strongly sequential updates should be integrated. Gradient or hybrid gates require task samples to compute gradient conflict, whereas GCWM's data-free gate uses geometry conflict available directly from task updates or checkpoints. Confidence intervals for the geometry and gradient target comparison and decompositions are in Appendices F.1 and F.5.

Figure 4. Geometry and gradient conflict are complementary. (a)--(c) reveal complementary failure modes: top-layer share is the fraction of top-ranked layers within each projection family, and (c) reports global step-level ∣ρs∣\vert \rho_s\vert .

4. Geometry Conflict Wasserstein Merging

We now turn the state-relative geometry findings in Sec. 3 into a data-free update-integration algorithm. Geometry-Conflict Wasserstein Merging (GCWM) operates on task vectors, estimates layer-wise geometry conflict, constructs a shared Wasserstein metric, and uses a conflict gate to control how strongly geometry-aware correction is applied. At continual step tt, GCWM then applies only the incremental change of the update, yielding a compatibility-controlled continual post-training merge. The state and global gaps summarize the evolving model state for analysis and are not directly optimized by GCWM. GCWM uses layer-wise geometry conflict over the complete history-aware active set to construct the shared Wasserstein metric and control the conflict gate.

4.1 Task Geometry and Conflict Gate

GCWM represents each active task update by its layer-wise covariance geometry. For an active update Δi\Delta_i and target linear layer ℓ\ell, let Δi(ℓ)∈Rdout×din\Delta_i^{(\ell)}\in\mathbb{R}^{d_{\mathrm{out}}\times d_{\mathrm{in}}}. We define

Ci(ℓ)=(Δi(ℓ))⊤Δi(ℓ)+λI,C_i^{(\ell)} = \big(\Delta_i^{(\ell)}\big)^\top \Delta_i^{(\ell)}+\lambda I,

which captures the dominant update subspaces and spectral energy while ensuring numerical stability.

To compare multiple active updates in a shared system, GCWM computes a truncated SVD

Δi(ℓ)≈Ui(ℓ)Σi(ℓ)(Vi(ℓ))⊤,\Delta_i^{(\ell)} \approx U_i^{(\ell)}\Sigma_i^{(\ell)}\big(V_i^{(\ell)}\big)^\top,

retains the principal right-singular directions, and forms

Q(ℓ)=orth ⁣([V1(ℓ),V2(ℓ),…,Vm(ℓ)]),Q^{(\ell)} = \mathrm{orth}\!\left( \left[ V_1^{(\ell)},V_2^{(\ell)},\dots,V_m^{(\ell)} \right] \right),

where mm is the number of active task updates. The projected geometry is

Bi(ℓ)=(Q(ℓ))⊤Ci(ℓ)Q(ℓ).B_i^{(\ell)} = \big(Q^{(\ell)}\big)^\top C_i^{(\ell)}Q^{(\ell)} .

The operators {Bi(ℓ)}i=1m\{B_i^{(\ell)}\}_{i=1}^m are used for conflict estimation and shared-metric construction.

For two projected geometries, GCWM defines layer-wise geometry conflict by the normalized Bures--Wasserstein discrepancy

γij(ℓ)=dB2 ⁣(Bi(ℓ),Bj(ℓ))tr ⁣(Bi(ℓ))+tr ⁣(Bj(ℓ))+ε,dB2(A,B)=tr(A)+tr(B)−2 tr ⁣((A1/2BA1/2)1/2),\gamma_{ij}^{(\ell)} = \frac{ d_{\mathrm{B}}^2\!\left(B_i^{(\ell)},B_j^{(\ell)}\right) }{ \mathrm{tr}\!\left(B_i^{(\ell)}\right) + \mathrm{tr}\!\left(B_j^{(\ell)}\right) + \varepsilon }, \qquad d_{\mathrm{B}}^2(A,B) = \mathrm{tr}(A)+\mathrm{tr}(B) - 2\,\mathrm{tr}\!\left( \left(A^{1/2}BA^{1/2}\right)^{1/2} \right),

where ε>0\varepsilon>0 is a stabilizer. Smaller γij(ℓ)\gamma_{ij}^{(\ell)} indicates more compatible task-induced geometries. GCWM aggregates pairwise conflicts and converts the result into a layer-wise gate:

g(ℓ)=∑i<jwijγij(ℓ),∑i<jwij=1,α(ℓ)=αmin⁡+(αmax⁡−αmin⁡) σ ⁣(κ(g(ℓ)−τ)).\begin{aligned} g^{(\ell)} &= \sum_{i<j}w_{ij}\gamma_{ij}^{(\ell)}, \qquad \sum_{i<j}w_{ij}=1, \\ \alpha^{(\ell)} &= \alpha_{\min} + (\alpha_{\max}-\alpha_{\min}) \,\sigma\!\big(\kappa(g^{(\ell)}-\tau)\big). \end{aligned}

Here wijw_{ij} are normalized task-pair weights, τ\tau is the conflict threshold, and κ\kappa controls gate sharpness. Thus, geometry conflict becomes an actionable layer-wise control signal rather than a purely diagnostic score.

4.2 Shared Wasserstein Metric and Gated Merge

Given {Bi(ℓ)}i=1m\{B_i^{(\ell)}\}_{i=1}^m, GCWM constructs a shared merging metric through the Gaussian Wasserstein barycenter

Bˉ(ℓ)=arg⁡min⁡B⪰0∑i=1mωidB2 ⁣(B,Bi(ℓ)),∑iωi=1.\bar{B}^{(\ell)} = \arg\min_{B\succeq 0} \sum_{i=1}^{m} \omega_i d_{\mathrm{B}}^2\!\left(B,B_i^{(\ell)}\right), \qquad \sum_i \omega_i=1 .

The barycenter Bˉ(ℓ)\bar{B}^{(\ell)} defines the local metric in which active updates are aligned before merging.

Let Δ^i(ℓ)=Δi(ℓ)Q(ℓ)\hat{\Delta}_i^{(\ell)}=\Delta_i^{(\ell)}Q^{(\ell)}. GCWM whitens the projected update, applies a base merge operator M\mathcal{M}, and recolors the result:

Δ~i(ℓ)=Δ^i(ℓ)(Bˉ(ℓ))−1/2,Δ~geo(ℓ)=M ⁣({Δ~i(ℓ)}i=1m;{ωi}i=1m),Δgeo(ℓ)=Δ~geo(ℓ)(Bˉ(ℓ))1/2(Q(ℓ))⊤.\begin{aligned} \tilde{\Delta}_i^{(\ell)} &= \hat{\Delta}_i^{(\ell)} \big(\bar{B}^{(\ell)}\big)^{-1/2}, \\ \tilde{\Delta}_{\mathrm{geo}}^{(\ell)} &= \mathcal{M}\!\left( \{\tilde{\Delta}_i^{(\ell)}\}_{i=1}^m; \{\omega_i\}_{i=1}^m \right), \nonumber \\ \Delta_{\mathrm{geo}}^{(\ell)} &= \tilde{\Delta}_{\mathrm{geo}}^{(\ell)} \big(\bar{B}^{(\ell)}\big)^{1/2} \big(Q^{(\ell)}\big)^\top . \end{aligned}

We instantiate M\mathcal{M} with weighted WUDI. The geometry-aware branch is then blended with an ungated plain merge:

Δmerge(ℓ)=α(ℓ)Δgeo(ℓ)+(1−α(ℓ))Δplain(ℓ).\Delta_{\mathrm{merge}}^{(\ell)} = \alpha^{(\ell)}\Delta_{\mathrm{geo}}^{(\ell)} + \big(1-\alpha^{(\ell)}\big)\Delta_{\mathrm{plain}}^{(\ell)} .

For clarity, Eqs. 8--9 present the projected form; the implementation uses the corresponding regularized full-space transform detailed in Appendix B.

4.3 Incremental Continual Update

GCWM is applied incrementally. At step tt, let At\mathcal{A}_t be the active set of task updates selected by the memory policy, containing the current update and optionally historical updates or the previous merged state. For each target layer, GCWM computes Δmerge,t(ℓ)\Delta_{\mathrm{merge},t}^{(\ell)} using Eqs. 4--10. Instead of reapplying the full merged update, GCWM applies only its change relative to the previous merged state:

Δinc,t(ℓ)=Δmerge,t(ℓ)−Δmerge,t−1(ℓ),Δmerge,0(ℓ)=0.\Delta_{\mathrm{inc},t}^{(\ell)} = \Delta_{\mathrm{merge},t}^{(\ell)} - \Delta_{\mathrm{merge},t-1}^{(\ell)}, \qquad \Delta_{\mathrm{merge},0}^{(\ell)}=0 .

The model update is θt(ℓ)=θt−1(ℓ)+ηtΔinc,t(ℓ),\theta_t^{(\ell)} = \theta_{t-1}^{(\ell)} + \eta_t\Delta_{\mathrm{inc},t}^{(\ell)}, where ηt\eta_t is a step coefficient. This rule keeps continual post-training tied to newly induced compatibility-controlled changes.

4.4 Theoretical Support for GCWM

We provide a fixed-step proposal analysis of GCWM relative to the plain merge. This analysis isolates the effect of geometry-aware correction before the incremental differencing in Eq. 11; the implementation applies the change between consecutive merged proposals. Under local smoothness, projected-geometry adequacy, and layer-wise metric-curvature assumptions stated in Appendix C, the relative loss effect of the GCWM correction is controlled by geometry conflict and metric displacement.

Let Θplain,t=θt−1+ηtΔplain,t,Θgcwm,t=θt−1+ηtΔmerge,t,\Theta_{\mathrm{plain},t} = \theta_{t-1}+\eta_t\Delta_{\mathrm{plain},t}, \qquad \Theta_{\mathrm{gcwm},t} = \theta_{t-1}+\eta_t\Delta_{\mathrm{merge},t}, where Δplain,t\Delta_{\mathrm{plain},t} and Δmerge,t\Delta_{\mathrm{merge},t} are the plain and GCWM merge proposals at step tt. Let B~t(ℓ)\widetilde B_t^{(\ell)} denote the implementation-aligned full-space metric induced by the shared Wasserstein metric.

Theorem 1. [Conflict-Controlled Integration]

Assume mt≥2m_t\ge2 and 0≤αmin⁡≤αmax⁡≤10\le \alpha_{\min}\le\alpha_{\max}\le1. Under the assumptions in Appendix C, the additional loss incurred on a previously acquired task uu by the GCWM proposal relative to the plain proposal satisfies

Lu(Θgcwm,t)−Lu(Θplain,t)≤ηt∑ℓcu,t(ℓ)gt(ℓ)+ηt22∑ℓdu,t(ℓ)∥Δmerge,t(ℓ)−Δplain,t(ℓ)∥B~t(ℓ)2,\mathcal L_u(\Theta_{\mathrm{gcwm},t}) - \mathcal L_u(\Theta_{\mathrm{plain},t}) \le \eta_t\sum_\ell c_{u,t}^{(\ell)}g_t^{(\ell)} + \frac{\eta_t^2}{2} \sum_\ell d_{u,t}^{(\ell)} \bigl\Vert \Delta_{\mathrm{merge},t}^{(\ell)} - \Delta_{\mathrm{plain},t}^{(\ell)} \bigr\Vert _{\widetilde B_t^{(\ell)}}^2,

where cu,t(ℓ),du,t(ℓ)≥0c_{u,t}^{(\ell)},d_{u,t}^{(\ell)}\ge0 are local constants.

Theorem 1 shows that the relative loss effect of the geometry-aware proposal is bounded by two quantities: the shared geometry conflict gt(ℓ)g_t^{(\ell)} and the metric displacement from the plain merge. The next result shows how the conflict gate controls this displacement.

Proposition 1. [Compatibility Regimes of Update Integration]

Assume 0≤αmin⁡≤αmax⁡≤10\le \alpha_{\min}\le\alpha_{\max}\le1. For each layer ℓ\ell, let

Dt(ℓ)=∥Δgeo,t(ℓ)−Δplain,t(ℓ)∥B~t(ℓ).D_t^{(\ell)} = \bigl\Vert \Delta_{\mathrm{geo},t}^{(\ell)} - \Delta_{\mathrm{plain},t}^{(\ell)} \bigr\Vert _{\widetilde B_t^{(\ell)}}.

Then ∥Δmerge,t(ℓ)−Δplain,t(ℓ)∥B~t(ℓ)=αt(ℓ)Dt(ℓ),∥Δmerge,t(ℓ)−Δgeo,t(ℓ)∥B~t(ℓ)=(1−αt(ℓ))Dt(ℓ).\bigl\Vert \Delta_{\mathrm{merge},t}^{(\ell)} - \Delta_{\mathrm{plain},t}^{(\ell)} \bigr\Vert _{\widetilde B_t^{(\ell)}} = \alpha_t^{(\ell)}D_t^{(\ell)}, \qquad \bigl\Vert \Delta_{\mathrm{merge},t}^{(\ell)} - \Delta_{\mathrm{geo},t}^{(\ell)} \bigr\Vert _{\widetilde B_t^{(\ell)}} = (1-\alpha_t^{(\ell)})D_t^{(\ell)}. Moreover, Eq. 6 implies gt(ℓ)≤τ⇒αt(ℓ)≤(αmin⁡+αmax⁡)/2g_t^{(\ell)}\le\tau \Rightarrow \alpha_t^{(\ell)}\le(\alpha_{\min}+\alpha_{\max})/2, whereas gt(ℓ)≥τ⇒αt(ℓ)≥(αmin⁡+αmax⁡)/2g_t^{(\ell)}\ge\tau \Rightarrow \alpha_t^{(\ell)}\ge(\alpha_{\min}+\alpha_{\max})/2. Thus, low-conflict layers receive weaker geometry-aware correction, while high-conflict layers receive stronger correction.

In summary, Theorem 1 and Proposition 1 show that GCWM is controlled by geometry conflict at both the loss and update levels. Proofs are provided in Appendices C and D.

[1]#1 [1]#1

Table 1. Domain-continual MMLU-Pro results on Qwen3-1.7B, 8B, and 14B. Scores are accuracies (%). Underlined MTL is a joint-training upper-bound reference; bold marks the best non-MTL result in each block. Data-free update-integration methods are shaded in blue.

MethodOverallBioBusChemCSEconEngHealthHistLawMathOtherPhilPhysPsych
Qwen3-1.7B
MTL44.465.151.643.047.352.337.141.930.223.353.235.839.546.452.9
Seq.\ SFT36.855.639.736.337.648.029.031.924.917.747.131.232.136.647.4
EWC40.058.244.639.640.048.632.335.824.117.755.431.731.542.946.6
FOREVER38.558.242.237.641.246.328.035.427.316.250.833.130.940.547.4
L\S41.161.448.239.944.048.834.238.827.420.649.832.936.443.249.5
AIMMerging41.860.749.741.445.950.836.636.924.418.456.434.432.342.347.5
OPCM41.762.749.040.444.849.734.539.327.520.550.733.136.843.850.4
GCWM43.564.951.042.246.651.736.141.129.021.952.734.838.645.752.3
Qwen3-8B
MTL65.383.370.665.966.374.454.365.458.040.276.158.260.767.573.9
Seq.\ SFT55.275.259.253.751.767.148.454.449.927.359.151.551.757.067.2
EWC60.478.866.263.061.071.151.859.449.929.972.352.955.363.268.2
FOREVER59.679.463.559.661.569.848.962.050.429.971.354.651.962.568.5
L\S62.479.468.367.065.470.552.661.551.732.377.355.552.166.067.4
AIMMerging62.978.169.167.965.171.353.062.452.032.178.755.551.967.467.9
OPCM61.978.866.862.462.870.551.461.954.938.172.055.157.563.970.0
GCWM63.781.468.964.264.772.752.763.756.438.874.356.659.165.872.2
Qwen3-14B
MTL68.686.274.072.467.877.056.867.361.939.979.464.661.772.277.6
Seq.\ SFT60.479.663.459.563.473.348.161.953.834.468.658.256.158.672.6
EWC65.384.270.067.063.474.152.666.359.333.580.961.255.969.171.7
FOREVER66.584.472.167.169.574.656.466.657.535.480.561.858.769.774.3
L\S65.682.770.869.264.873.754.164.359.137.776.061.758.969.174.3
AIMMerging66.483.971.770.165.674.754.665.059.737.777.162.459.570.075.3
OPCM66.683.771.870.265.874.755.165.360.138.777.062.759.970.175.3
GCWM67.883.873.172.372.776.255.866.159.636.183.461.559.573.372.3

5. Experiments

We evaluate GCWM as a data-free update-integration method under domain and capability shifts. Our main comparisons focus on data-free merging baselines; sequential, regularized, and replay-based methods are included as reference continual-training pipelines.

5.1 Setup

Models: We use Qwen3 backbones at 0.6B, 1.7B, 4B, 8B, and 14B. \ Training data: For domain-continual training, we use and form a 14-task sequence with 1k samples per sub-domain. For capability-continual training, we use 30k math samples from and 30k code samples from. More setup details are provided in Appendix G.1 and G.2. \ Baselines: Our main baselines are data-free update-integration methods, including Localize-and-Stitch, AIMMerging, and OPCM. Seq.\ SFT, EWC, and FOREVER are reported as reference continual-training pipelines because they use additional regularization or replay. Non-continual merging like TA, TIES, DARE are also reported in Appendix G.4. \ Benchmarks: Domain-continual performance is evaluated on the 14 MMLU-Pro sub-categories. Capability-continual performance is evaluated on GSM8K, MATH-500, MBPP, HumanEval, GPQA-Diamond, and MMLU-Pro. \ Remark 1: All reported performance scores are averaged over five independent evaluation runs. \ Remark 2: The evaluation code we employ strictly adheres to the Qwen3 Technical Report

5.2 Domain-Continual Post-Training

We first evaluate domain-continual post-training on the 14-domain MMLU-Pro sequence. This setting tests whether GCWM can integrate domain-specific updates without accessing training data during the merge stage. Table 1 reports full results on Qwen3-1.7B, 8B, and 14B. MTL is included as a joint-training upper-bound reference, while the main comparison is among data-free update-integration methods. GCWM gives the strongest non-MTL overall performance at all three scales. It improves over the best data-free baseline by +1.61+1.61, +0.74+0.74, and +1.23+1.23 points on Qwen3-1.7B, 8B, and 14B, respectively. The gains are broad rather than driven by a single domain: GCWM improves over AIMMerging on 12/14 domains at 1.7B, 10/14 domains at 8B, and 9/14 domains at 14B. These results support the role of geometry conflict as a practical control signal for data-free continual update integration. Complete results for all model scales are reported in Appendix G.5.

5.3 Capability-Continual Post-Training

We next evaluate capability-continual post-training with sequential math and code updates. Table 2 reports the performance of Qwen3-1.7B and 14B across three key domains: knowledge (GPQA-Diamond, MMLU-Pro), math (GSM8K, MATH-500), and code (HumanEval, MBPP). At 1.7B, GCWM gives the best data-free average (58.3), improving over the strongest data-free baseline (OPCM, 56.8) by +1.52 points and leading on all six benchmarks. At 14B, GCWM remains the strongest data-free method on average (74.3 vs. 72.9 for OPCM) and leads on GPQA-Diamond, GSM8K, HumanEval, and MMLU-Pro. Benchmark-specific gaps to MTL can reflect capability acquisition, data mixing, and update integration, so the GSM8K--MATH-500 contrast does not isolate forgetting or imply that deeper reasoning is intrinsically more forgettable. FOREVER can be higher in some settings because it revisits data through replay and additional sequential optimization; we include it as a replay-based reference, while the primary comparison is among data-free update-integration methods. These results show that geometry-conflict-controlled integration extends beyond domain transfer to heterogeneous capability updates. Full five-scale results are provided in Appendix G.6.

Table 2. Capability-continual results on Qwen3-1.7B and 14B. Scores are accuracies or pass@1 (%). MTL is a joint-training reference; Seq.\ SFT, EWC, and FOREVER are training-pipeline references. Bold marks the best data-free method.

MethodAvg.GPQA-D.GSM8KHumanEvalMATH-500MBPPMMLU-Pro
1.7B14B1.7B14B1.7B14B1.7B14B1.7B14B1.7B14B1.7B14B
MTL57.174.626.733.376.192.161.684.263.687.856.880.557.569.8
Seq.\ SFT51.970.418.243.470.695.864.086.664.266.259.163.435.367.2
EWC54.773.524.843.475.795.461.686.067.868.057.678.640.569.4
FOREVER58.375.927.354.076.796.462.287.269.669.266.975.947.372.8
L\S52.471.321.338.471.094.162.883.757.576.558.478.543.556.8
AIMMerging53.472.221.738.872.495.264.084.758.677.459.579.444.357.5
OPCM56.872.923.238.073.094.565.280.567.679.759.978.651.966.3
GCWM58.374.326.339.979.095.867.486.663.478.261.576.752.068.8

5.4 Ablations and Analysis

We ablate two merge-time components of GCWM: the conflict gate and the shared Wasserstein metric. All variants use the same task experts and evaluation protocol within each setting, differing only in the integration rule. The w/o gate variant removes conflict-conditioned gating and applies the geometry-aware branch uniformly, while w/o Wasserstein barycenter replaces the shared Wasserstein barycenter with a mean covariance metric. Fig. 5 shows that full GCWM improves average performance over w/o gate and w/o Wasserstein barycenter by 1.58 and 1.31 points on Qwen3-1.7B and 4.61 and 3.73 points on Qwen3-8B, respectively. The benchmark-level trade-offs show that compatibility-controlled integration improves aggregate performance without uniformly improving every capability. These results support the gate as a layer-wise control signal and the Wasserstein metric as a shared geometry for data-free update integration. Domain-continual ablations on Qwen3-0.6B and full capability breakdowns are reported in Appendix H.

Figure 5. Capability-continual GCWM ablations on Qwen3-1.7B and 8B. All variants use the same task experts within each scale; WB denotes Wasserstein barycenter.

Remark 3: We also profile GCWM runtime and memory on Qwen3-8B and 14B in Appendix I, since these costs are the practical bottleneck for data-free update integration. \ Remark 4: Appendix J reports Qwen3-8B rank and gate-parameter sensitivity.

6. Conclusion

We studied LLM continual post-training through task geometry, analyzing update norm, subspace alignment, gradient conflict, and geometry conflict across model scales and continual strategies. Our main finding is that forgetting is a state-relative update-integration failure: harmful steps occur when task-induced covariance geometries become incompatible with the geometry of the evolving model state. This explains why raw drift and isolated pairwise compatibility are insufficient, and why geometry conflict serves as both an explanatory signal for forgetting and a control signal for sequential update integration. Geometry-Conflict Wasserstein Merging (GCWM) operationalizes this insight by constructing a Wasserstein shared metric from task-induced covariance geometry and gating data-free update integration by geometry conflict. Across domain-continual and capability-continual settings, GCWM improves retention and final performance over data-free baselines without replay data.

7. Limitations

Our analysis and experiments focus on Qwen3-scale open LLMs and on domain and capability continual post-training tasks built from public reasoning, knowledge, math, and code benchmarks. Although the state-relative geometry signal is consistent across scales and methods, it should be viewed as an explanatory and control signal rather than a proof of causal necessity for all forms of forgetting. GCWM is data-free at merge time and is therefore most relevant when historical data are unavailable or replay is undesirable; replay-based training can still be stronger when it is allowed to repeatedly revisit past data. The method also assumes access to task-specific updates or checkpoints and incurs additional CPU cost for geometry construction and Wasserstein metric computation, though this cost is offline and does not affect inference.

This work was supported by the Hong Kong RGC (TRS: T41-517/25-N; GRF: 15228325), ITC RAISe+ (RAI/24/1/086A), and the Research Institute for Generative Artificial Intelligence at PolyU.

BibTeX

@misc{wang-2026-gcwm,
      title={Geometry Conflict: Explaining and Controlling Forgetting in LLM Continual Post-Training},
      author={Yuanyi Wang and Yifan Yang and Su Lu and Yanggan Gu and Pengkai Wang and Wenjun Wang and Zhaoyi Yan and Congkai Xie and Jianmin Wu and Jialun Cao and Shing-Chi Cheung and Hongxia Yang},
      year={2026},
      url={https://github.com/wyy-code/GCWM},
}