LIF \(M!\) Ordering Space — Research CheckpointLIF \(M!\) Ordering Space — 目前研究 Checkpoint
A mathematical, system-level explanation of the current architecture: from JEPA input and continuous LIF geometry to factorial ordering states, functional prediction, adjacent updates, and the remaining addressing bottleneck.
從 JEPA 輸入、continuous LIF 幾何、\(M!\) ordering state、functional prediction、adjacent update,到目前真正 addressing bottleneck 的完整數學架構說明。
§1Current system in one picture目前系統一張圖看完
The project is no longer one monolithic “SNN model.” It is best understood as four mathematically distinct modules: representation, functional evaluation, action selection, and temporal actuation. The main progress came from separating these modules so that one component cannot silently solve the job of another.
現在這個研究已經不是一個混在一起的「SNN model」。最乾淨的理解方式,是把它拆成四個數學上不同的模組:representation、functional evaluation、action selection、temporal actuation。真正的重要進展,就是把這四件事拆開,避免某個大模型偷偷把其他模組的工作全部做掉。
Current strongest separation: the factorial state substrate and the M6 LIF actuator are real and well controlled. The unresolved problem is the functional address decoder: how to infer which local edge is useful without an \(M!\)-sized key or a generic universal learner.
目前最重要的分離:\(M!\) state substrate 與 M6 的 LIF actuator 已經是很扎實的部分。真正還沒解掉的是 functional address decoder:怎麼不用 \(M!\)-size key、也不用 generic 萬能模型,就知道「現在到底該走哪一條 local edge」。
§2Symbols and state spaces符號與 state space
§3Training data and experimental unitsTraining data 與實驗單位
The common formal task is the frozen Moving-Shapes JEPA line. The selected functional/navigation lineage uses 100,000 train, 10,000 validation, and 10,000 test samples, with formal seeds 0/1/2. M6 is the primary mechanism-discovery size.
目前共用的 formal task 是 frozen Moving-Shapes JEPA lineage。selected functional / navigation 線使用 train 100,000、validation 10,000、test 10,000,formal seeds 為 0/1/2;M6 是主要 mechanism-discovery 尺度。
| Data object資料物件 | Meaning意義 | Used for用途 |
|---|---|---|
task sample x | Moving-Shapes context/target sampleMoving-Shapes context / target sample | JEPA prediction |
(x, π) | same sample evaluated at a specified ordering state同一 sample 在指定 ordering state 下的 counterfactual | \(L(x,\pi)\) |
(x, π, k) | one adjacent candidate edge一條 adjacent candidate edge | \(U_k\), ranking labels |
(y,\hat y,e,\Delta e) | JEPA residual decomposition recordJEPA residual decomposition record | residual/action-effect mechanism auditresidual / action-effect mechanism audit |
One true evaluator call yields \(L(x,\pi)\). Utilities and pairwise ranking labels can then be derived algebraically from already-evaluated candidate losses; they are not new independent task observations. This distinction prevents factorial supervision from becoming a hidden key.
一次真正 evaluator call 才產生 \(L(x,\pi)\)。utilities、pairwise ranking labels 可以由已經算過的 candidate losses 代數推導,它們不是新的 independent task observations。這個 distinction 用來防止 factorial supervision 偷偷變成 hidden key。
§4JEPA front end: prediction defines functionJEPA 前端:prediction objective 定義「有用」
Permutation states are not intrinsically good or bad. JEPA supplies the task objective. Conceptually:
permutation states 本身沒有「好或壞」。真正讓它有 functional meaning 的是 JEPA prediction objective:
The exact frozen loss \(\ell\) is authoritative. If the implementation contains normalization, cosine geometry, or auxiliary terms, it must not be silently replaced by MSE.
真正的 frozen loss \(\ell\) 必須以 implementation 為準。如果實際包含 normalization、cosine geometry 或其他項目,就不能擅自當成 MSE。
Established: ordering has predictive signal, and the frozen task induces a real counterfactual functional landscape over permutation states.
已成立:ordering 有 predictive signal,而且 frozen task 確實在 permutation states 上形成真實 counterfactual functional landscape。
§5Continuous score geometry before spikesSpike 前的 continuous score geometry
Ordering is invariant to a common score shift, so the intrinsic score geometry has at most \(M-1\) degrees of freedom.
所有 score 一起加 common shift 不會改 ordering,所以 intrinsic score geometry 最多只有 \(M-1\) 個自由度。
The mechanism-first line prefers a fixed/constrained map rather than a generic MLP:
mechanism-first 路線偏好 fixed / constrained map,而不是 generic MLP:
For M6, \(u\in\mathbb R^5\) and \(W\in\mathbb R^{6\times5}\). Hyperplanes \(s_i=s_j\) divide the centered 5-D space into ordering chambers. All 720 M6 permutations are reachable at \(d=5\).
M6 時 \(u\in\mathbb R^5\)、\(W\in\mathbb R^{6\times5}\)。超平面 \(s_i=s_j\) 把 centered 5-D space 切成 ordering chambers;\(d=5\) 時全部 720 permutations 都可 reach。
M6 has \(\binom62=15\) pair margins, but they are redundant: \(D_{12}+D_{23}=D_{13}\). The intrinsic dimension remains five.
M6 雖然有 \(\binom62=15\) 個 pair margins,但它們是 redundant,例如 \(D_{12}+D_{23}=D_{13}\)。intrinsic dimension 仍然是 5。
§6LIF timing mapLIF timing map
In the valid regime, \(t^*\) decreases monotonically with current:
在有效區間內,\(t^*\) 對 current 單調遞減:
Therefore shared monotone LIF physically realizes score ordering in time. Exact/event timing matters because coarse discrete timesteps collapse distinct crossings into ties.
所以 shared monotone LIF 的角色,是把 score ordering 物理化到時間軸。exact/event timing 很重要,因為粗 timestep 會把其實不同的 crossings 壓成 ties。
The current clean role of LIF is temporal realization and local actuation. A specifically LIF-only functional advantage has not yet been established.
目前 LIF 最乾淨的角色是 temporal realization + local actuation;「只有 LIF 才有的額外 functional advantage」尚未成立。
§7Permutation state and factorial capacityPermutation state 與 factorial capacity
Nominal capacity is not the same as reachable, used, effective, controllable, or functional capacity. Ordering cannot create information:
nominal capacity 不等於 reachable / used / effective / controllable / functional capacity。ordering 不能創造資訊:
The possible value of \(M!\) is combinatorial organization, not free information.
\(M!\) 可能帶來的是 combinatorial organization,不是免費資訊。
§8Mathematical adjacent-swap LIF actuator數學化 adjacent-swap LIF actuator
Let \(u=\pi_k\), \(v=\pi_{k+1}\) and currently \(I_u>I_v\). Take:
令 \(u=\pi_k\)、\(v=\pi_{k+1}\),目前 \(I_u>I_v\)。取:
Choose \(\epsilon\) within neighbor and spike-feasibility margins. Then only the target pair reverses and shared monotone LIF gives:
只要 \(\epsilon\) 落在 neighbor / spike-feasibility margins 內,就只會翻轉 target pair:
The analytic actuator passed all 3600 directed edges. The minimal learned actuator selected five rank-conditioned scalar margins. After target-independent stabilization, the ordered source-target path audit passed:
analytic actuator 通過全部 3600 directed edges;minimal learned actuator 只需要 5 個 rank-conditioned scalar margins。加上 target-independent stabilization 後:
Interpretation: at M6, physical execution of a correct local move is no longer the main bottleneck.
解讀:至少在 M6,「正確 local move 能不能被 LIF 真的執行」已經不是主要瓶頸。
§9Task-induced functional fieldTask 誘導出的 functional field
This exact potential structure is why learned scalar potentials are attractive: \(\hat U(A\to B)=V(A)-V(B)\) guarantees antisymmetry and zero circulation.
這也是 scalar potential 很有吸引力的原因:若 \(\hat U(A\to B)=V(A)-V(B)\),antisymmetry 與 zero circulation 自動成立。
Functional headroom remains strongfunctional headroom 並沒有隨 M 消失
| M | normalized oracle headroom | beneficial-state rate |
|---|---|---|
| 6 | 1.2328 | 0.9130 |
| 8 | 1.2407 | 0.9501 |
| 10 | 1.2378 | 0.9773 |
| 12 | 1.3228 | 0.9853 |
At M12, about 98.5% of sampled states still have at least one beneficial adjacent move. The issue is not absence of useful actions.
M12 sampled states 裡約 98.5% 仍至少有一條 beneficial adjacent move,所以問題不是「沒有 useful action」。
§10What one-step ranking has already shownOne-step ranking 已經證明什麼
The M6 potential-ranked controller is useful as a control: compositional state information contains real one-step action signal.
M6 potential-ranked controller 現在最適合當 control:它證明 compositional state information 的確含有真實 one-step action signal。
The selected controller used 6,913 parameters with hidden width 64, without a 720-way head or factorial lookup.
selected controller 是 6,913 parameters、hidden width 64,沒有 720-way head 或 factorial lookup。
| metric | value |
|---|---|
| mean gain | 0.02219 |
| navigation efficiency \(\eta\) | 0.5703 |
| harmful rate | 0.1080 |
| beneficial precision | 0.8201 |
| listwise top-1 | 0.6530 |
| NDCG | 0.9012 |
The important result is one-step learnability, not final closed-loop success. Once the policy updates \(\pi\), it changes its own input distribution; a first-step success does not guarantee recursive safety.
真正重要的是 one-step learnability,不是把這個 controller 當 final closed-loop success。一旦 policy update \(\pi\),它就改變自己的 input distribution;第一步有效不代表後面 recursive safety 也成立。
§11Current mechanistic direction: JEPA residual action-effect factorization目前機制方向:JEPA residual action-effect factorization
Instead of directly learning scalar \(U_k\), expose the current JEPA error and the vector effect of one action:
目前改成不要直接硬學 scalar \(U_k\),而是把 JEPA 現在錯在哪裡、以及一個 action 對 error vector 做了什麼拆開:
If the verified frozen loss is sum squared error:
如果 verified frozen loss 是 sum squared error:
For mean-MSE:
若是 mean-MSE:
This explains why the same action can flip from helpful to harmful across contexts: \(\Delta e_k\) may be similar while the current residual \(e\) changes direction. Scalar globality therefore does not automatically imply a globally complex action mechanism.
這可以解釋為什麼同一 action 在不同 context 會從 helpful 翻成 harmful:\(\Delta e_k\) 可能很類似,但 current residual \(e\) 的方向不同。所以 scalar globality 不代表 action mechanism 必然也 global。
Evidence boundary: you stated that the 2.7 formal run and strict audit are complete, but its metrics/flags were not provided in this source set. This report therefore documents the 2.7 mathematics without inventing the formal scientific result.
證據邊界:你已說 2.7 formal 與 strict audit 完成,但目前提供給我的內容沒有 2.7 實際 metrics / flags;所以這份報告會寫完整 2.7 數學架構,但不自行猜它的 formal scientific result。
§12Full mathematical loop: input → prediction → update → prediction → update完整數學迴圈:input → prediction → update → 再 prediction → 再 update
Step 1 · task sampletask sample
\[x=(x_{\rm context},x_{\rm target}).\]Step 2 · task representationstask representations
\[z=E_{\rm ctx}(x_{\rm context}),\qquad y=E_{\rm tgt}(x_{\rm target}).\]Step 3 · open ordering geometry打開 ordering geometry
\[u=C(z),\qquad s=W_{\rm orth}u,\] \[W_{\rm orth}^\top W_{\rm orth}=I,\qquad 1^\top W_{\rm orth}=0.\]The exact frozen \(C\) is implementation-defined; the mechanism rule is that this coordinate map cannot become an unconstrained factorial-memory network.
實際 frozen \(C\) 以 implementation 為準;mechanism rule 是這個 coordinate map 不能變成 unconstrained factorial-memory network。
Step 4 · relational staterelational state
\[D_{ij}=s_i-s_j,\qquad P_{ij}=\operatorname{sign}(D_{ij}).\]Step 5 · real LIF timingreal LIF timing
\[s/I\longrightarrow t_i^*=-\tau\ln\left(1-\frac{V_{\rm th}}{RI_i}\right),\] \[\pi_r=\operatorname{argsort}(t_1^*,\ldots,t_M^*).\]Step 6 · prediction at current state目前 state 下 prediction
\[\hat y_r=F_{\rm pred}(z,\mathcal R(\pi_r)),\qquad L_r=\ell(y,\hat y_r).\]Step 7 · local candidate fieldlocal candidate field
\[\pi_r^{(k)}=\operatorname{swap}_k(\pi_r),\qquad L_{r,k}=L(x,\pi_r^{(k)}),\] \[U_{r,k}=L_r-L_{r,k},\qquad k=1,\ldots,M-1.\]Step 8 · decisiondecision
\[a_r^*=\arg\max(0,U_{r,1},\ldots,U_{r,M-1}).\]Autonomous operation replaces true \(U\) with a compact estimate \(\hat U\). This is the unresolved address decoder.
Autonomous operation 需要用 compact estimate \(\hat U\) 取代真實 \(U\)。這就是目前尚未解掉的 address decoder。
Step 9 · physical updatephysical update
For \(a_r=k\), \(u=\pi_{r,k}\), \(v=\pi_{r,k+1}\):
若 \(a_r=k\),\(u=\pi_{r,k}\)、\(v=\pi_{r,k+1}\):
\[\mu_r=\frac{I_u+I_v}{2},\qquad I'_u=\mu_r-\epsilon_k,\qquad I'_v=\mu_r+\epsilon_k.\] \[\pi_{r+1}=\operatorname{swap}_k(\pi_r)\]after the target-independent stabilization and exact LIF timing.
在 target-independent stabilization + exact LIF timing 後成立。
Step 10 · predict again再 prediction 一次
\[\hat y_{r+1}=F_{\rm pred}(z,\mathcal R(\pi_{r+1})),\qquad L_{r+1}=\ell(y,\hat y_{r+1}).\] \[\pi_{r+1}\rightarrow U_{r+1,k}\rightarrow a_{r+1}\rightarrow\pi_{r+2}\rightarrow L_{r+2}\rightarrow\cdots\]This is the exact distinction between one-step actionability and closed-loop navigability: the policy changes the state it will see next.
這就是 one-step actionability 跟 closed-loop navigability 的精確差別:policy 會改變它下一步自己要看到的 state。
§13What is trained, and with what loss?到底哪些東西有 training?training data / loss 是什麼?
A · JEPA / predictor training
The predictor is trained on Moving-Shapes context/target prediction and then frozen for counterfactual landscape studies.
predictor 先在 Moving-Shapes context/target prediction 上 training,之後 functional landscape 系列都 freeze。
\[\mathcal L_{\rm JEPA}=E_x[\ell(y(x),\hat y(x))].\]Formal selected split: 100k train / 10k validation / 10k test. The exact \(\ell\) must follow source implementation.
selected formal split:100k train / 10k validation / 10k test。實際 \(\ell\) 必須跟 source implementation 一致。
B · Counterfactual functional labels
\[L_0=L(x,\pi),\qquad L_k=L(x,\operatorname{swap}_k\pi),\] \[U_k=L_0-L_k.\]There is no supervised permutation class. Ranking labels are derived from candidate losses:
沒有 supervised permutation class。ranking labels 是由 candidate losses 導出:
\[y_{ab}=\operatorname{sign}(L_b-L_a).\]C · Potential ranking losses
A representative pairwise objective is:
代表性的 pairwise objective:
\[\mathcal L_{\rm pair}=\sum_{aListwise versions rank all actions jointly; harm/regret-aware variants put extra cost on selecting truly harmful candidates. Validation, not test, selects temperature/weights.listwise 版本一次 rank 全部 actions;harm/regret-aware 會對真正 harmful candidate 額外加權。temperature / weights 用 validation 選,不用 test。
D · LIF actuator learning
The actuator never learns task utility. It receives a desired adjacent edge and uses only structural swap supervision. M6's selected learned actuator has five rank-conditioned scalar margins:
actuator 完全不學 task utility;它收到 desired adjacent edge,只做 structural swap supervision。M6 selected learned actuator 只有五個 rank-conditioned scalar margins:
\[\epsilon=\epsilon_k,\qquad k=1,\ldots,5.\]E · JEPA residual-effect training target
If the residual route requires a learned effect model, its primary target is the vector \(\Delta e_k\), not scalar utility:
如果 JEPA residual 路線需要 learned effect model,primary target 是 vector \(\Delta e_k\),不是 scalar utility:
\[\mathcal L_{\Delta e}=E\|\widehat{\Delta e}_k-\Delta e_k\|_2^2.\]Utility is then reconstructed by the verified loss geometry:
utility 再由 verified loss geometry 重建:
\[\hat U_k=F_\ell(e,\widehat{\Delta e}_k).\]§14Complexity and the anti-universal-key rule複雜度與 anti-universal-key rule
| Operation操作 | complexity | Meaning意義 |
|---|---|---|
| score state | \(O(M)\) | M continuous values |
| pair differences | \(O(M^2)\) | \(\binom M2\) relational margins |
| hard decode | \(O(M\log M)\) | sort spike times排序 spike times |
| candidate actions | \(M-1\) | adjacent swaps |
| one analytic swap | \(O(1)\) local edit | target pair current edittarget pair current edit |
| nominal states | \(M!\) | implicit chambers, not output neuronsimplicit chambers,不是 output neurons |
- forbid\(O(M!)\) permutation embeddings / logits / lookup tables
- forbidM>6 factorial training enumerationM>6 factorial training enumeration
- forbidlarge generic controller as the real mechanismlarge generic controller 成為真正 mechanism
- watchdense polynomial/global structures that are polynomial but scientifically non-compactdense polynomial / global structures,雖然 polynomial 但可能科學上仍不 compact
§15The remaining bottleneck, mathematically現在真正瓶頸,用數學講清楚
- answeredCan the state exist?state 能不能存在?
\(M=6,d=5\Rightarrow720/720\). - answeredCan LIF execute a requested local move?LIF 能不能執行指定 local move?
\(3600/3600\) adjacent edges; stabilized \(518400/518400\) paths. - openCan a compact mechanism know which move is functionally correct?小型 mechanism 能不能知道哪一條 move 對 task 是對的?
Local action does not imply local decisionLocal action 不代表 local decision
| M | pair relations \(N\) | q90 dependency | q90/N |
|---|---|---|---|
| 6 | 15 | 6* | 0.40 |
| 8 | 28 | 15 | 0.54 |
| 10 | 45 | 28 | 0.62 |
| 12 | 66 | 45 | 0.68 |
* M6 q90 used a conservative repair; M8–M12 are the stronger evidence for the globalizing trend.
* M6 q90 經過保守 repair;globalizing trend 主要以 M8–M12 為較強證據。
The unresolved source question is whether this globality is intrinsic to the task, introduced by ordering, or produced by collapsing a vector prediction error into scalar utility. A direct coordinate-linear instrument had held-out \(R^2\approx0\) and sign accuracy near 0.5, so it did not answer that question.
還沒回答的是:globality 是 task 本來就 global、ordering 造成,還是把 vector prediction error 壓成 scalar utility 後產生?direct coordinate-linear instrument 的 held-out \(R^2\approx0\)、sign accuracy 近 0.5,所以沒有回答成功。
The JEPA residual route therefore asks whether a global-looking scalar \(U_k\) factors into a current error state \(e\) and a simpler action-effect vector \(\Delta e_k\).
因此 JEPA residual 路線改問:看起來很 global 的 scalar \(U_k\),是否其實可以拆成 current error state \(e\) 與更簡單的 action-effect vector \(\Delta e_k\)。
§16Evidence that matters to the current architecture真正會影響目前架構的核心證據
| Claim命題 | Evidence證據 | Status |
|---|---|---|
| M6 full ordering geometryM6 full ordering geometry | \(d=5\Rightarrow720/720\) | established |
| M6 local actuationM6 local actuation | \(3600/3600\) | established |
| M6 stabilized path actuationM6 stabilized path actuation | \(518400/518400\) | established |
| Task-induced functional landscapetask-induced functional landscape | oracle headroom remains strong M6–M12oracle headroom 在 M6–M12 保持強 | established |
| Oracle action executed by real LIForacle action 被 real LIF 執行 | exact execution 1.0 | established |
| One-step ranking signalone-step ranking signal | \(\eta\approx0.57\), harmful \(\approx0.108\) | established |
| Compact direct pair/rank utility lawcompact direct pair/rank utility law | insufficient不足 | not established |
| Bounded local/sparse/low-rank dependency at M8–M12M8–M12 bounded local/sparse/low-rank dependency | not supported未支持 | not established |
| Globality source: task vs orderingglobality 來源:task vs ordering | measurement instrument insufficientmeasurement instrument 不足 | open |
| Compact JEPA residual action-effect lawcompact JEPA residual action-effect law | 2.7 metrics not supplied here此處尚未提供 2.7 metrics | pending evidence |
§17What the architecture is now所以現在的架構到底是什麼
- resolvedState substrate: centered continuous score geometry → first-spike permutation.State substrate:centered continuous score geometry → first-spike permutation。
- resolvedState write/update: adjacent mathematical current edit → stabilized real LIF actuation.State write/update:adjacent mathematical current edit → stabilized real LIF actuation。
- resolvedTask value: frozen JEPA/task predictor defines \(L(x,\pi)\); utility is an exact loss difference.Task value:frozen JEPA/task predictor 定義 \(L(x,\pi)\);utility 是 exact loss difference。
- partialOne-step address decoder: real signal exists, but the compact direct law is not final.One-step address decoder:確實有 signal,但 compact direct law 還不是 final。
- openRecursive address decoder: must remain safe after its own updates change the state distribution.Recursive address decoder:自己的 update 改變 state distribution 後仍要保持安全。
- currentCurrent mechanistic hypothesis: \(e+\Delta e_k+\) known JEPA loss geometry.目前 mechanistic hypothesis:\(e+\Delta e_k+\) known JEPA loss geometry。
What changed conceptually一路上真正改掉的是什麼
§18Small conclusion目前的小結論
1. A small continuous LIF score geometry can support a factorial number of discrete first-spike ordering states. At M6, five intrinsic score dimensions are enough for all 720 permutations.
1. 少量 continuous LIF score geometry 可以支撐 factorial 數量的 first-spike ordering states。M6 只需要 5 個 intrinsic score dimensions,就能形成完整 720 permutations。
2. The state space is physically writable: a tiny mathematical LIF actuator executes adjacent moves, and stabilized M6 path control is complete.
2. 這個 state space 是真的可寫入:很小的 mathematical LIF actuator 可以執行 adjacent moves,而且 stabilized M6 path control 已完整通過。
3. The frozen JEPA/task objective gives the ordering graph real functional structure. Useful local moves remain abundant up to M12.
3. frozen JEPA/task objective 確實讓 ordering graph 有真實 functional structure;useful local moves 到 M12 仍然很多。
4. The hard part is functional addressing: deciding which edge is useful with a compact scalable law, not generating the states or executing M6 swaps.
4. 現在最難的是 functional addressing:如何用 compact scalable law 判斷哪條 edge 有用;不是產生 states,也不是執行 M6 swaps。
5. The current mechanistic question is whether scalar edge utility is complicated only because it combines a current JEPA residual \(e\) with an action correction \(\Delta e_k\). If \(\Delta e_k\) is compact, global decision context does not imply a global action mechanism.
5. 現在最值得問的是:scalar edge utility 看起來複雜,會不會只是因為它把 current JEPA residual \(e\) 和 action correction \(\Delta e_k\) 組合在一起?如果 \(\Delta e_k\) 很 compact,那 decision 需要 global context 並不代表 action mechanism 也 global。
One-sentence mental model: a compact factorial state substrate and a reliable local write mechanism already exist; the open question is whether the task admits a compact read/address law, possibly in JEPA residual geometry rather than permutation identity alone.
一句話心智模型:「compact factorial state substrate + reliable local write mechanism」已經存在;現在真正未知的是 task 有沒有 compact read/address law,而這個 law 可能要從 JEPA residual geometry 找,不一定只靠 permutation identity。