Gary's LIF Ordering SystemGary's LIF Ordering System
Factorial first-spike states, the 5-D Navigator, exact gradient addressing, learnable LIF parameters, the 3.0–3.3 series, and a careful comparison with snnTorch.Factorial first-spike states、5-D Navigator、exact gradient addressing、learnable LIF parameters、3.0–3.3 系列,以及與 snnTorch LIF 的完整差異。
§1Definition: what is actually different in this LIF system?定義:這個 LIF 系統真正不同的是什麼?
The base LIF differential equation is standard. The system-level formulation is the distinctive part: use first-spike ordering as the state, treat the resulting permutation as an element of \(S_M\), and use an \(M-1\)-dimensional Navigator to gradient-address a small set of LIF timing controls.
基本 LIF 微分方程不是新的。真正不同的是系統層做法:把 first-spike ordering 本身當 state,讓 permutation 成為 \(S_M\) 的元素,再用 \(M-1\) 維 Navigator 對少量 LIF timing controls 做 gradient addressing。
For M6 the smallest formally validated variants use only six adjustable values per run, while the Navigator has five dimensions. No 720-way classifier or 720-entry lookup is required.
M6 目前最小且正式通過的版本,每次 run 只需要 6 個可調數字,而 Navigator 只有 5 維;沒有 720-way classifier,也沒有 720-entry lookup。
§2The continuous LIF equation used in 3.0–3.23.0–3.2 使用的 continuous LIF 方程
For the clean first-spike experiments, set \(V_i(0)=0\), take \(V_{rest}=0\), and hold \(I_i\) constant before the first spike.
乾淨的 first-spike 實驗先令 \(V_i(0)=0\)、\(V_{rest}=0\),並在 first spike 前令 \(I_i\) 為常數。
Because the representation uses only the first spike, post-spike reset is irrelevant to the 3.0–3.2 ordering definition.
因為 representation 只讀第一個 spike,所以 3.0–3.2 的 ordering 定義不依賴 first spike 後的 reset。
§3Every variable in the function函式裡每一個東西的意思
| symbol | meaning意思 | if increased變大時 |
|---|---|---|
| \(V_i(t)\) | membrane potential; dynamic statemembrane potential;動態 state | — |
| \(I_i\) | input current / external driveinput current / 外部 drive | \(t_i^*\downarrow\) |
| \(\tau_i\) | membrane time constantmembrane time constant;膜時間尺度 | \(t_i^*\uparrow\) |
| \(V_{{th},i}\) | spike thresholdspike threshold;發射門檻 | \(t_i^*\uparrow\) |
| \(R_i\) | membrane resistance; current-to-voltage gainmembrane resistance;current→voltage 增益 | \(t_i^*\downarrow\) |
| \(t_i^*\) | first-spike time第一次碰 threshold 的時間 | smaller = earlier越小越早 |
Notation warning: here \(R_i\) means membrane resistance. In the snnTorch documentation, the symbol \(R\) inside the discrete recurrence denotes a reset indicator/mechanism, not resistance.
符號警告:本文的 \(R_i\) 是 membrane resistance。snnTorch 官方離散 recurrence 裡也寫 \(R\),但那個 \(R\) 是 reset indicator / mechanism,不是 resistance。
§4Exact first-spike solution and exact gradientsExact first-spike 解與 exact gradients
This is why 3.0–3.2 can differentiate first-spike time directly instead of differentiating a hard binary spike.
這就是 3.0–3.2 可以直接對 first-spike time 微分,而不需要對 hard binary spike 微分的原因。
| parameter | exact derivative | sign |
|---|---|---|
| \(I\) | \(\displaystyle \frac{\partial t^*}{\partial I}=-\frac{\tau V_{th}}{I(RI-V_{th})}\) | negative |
| \(\tau\) | \(\displaystyle \frac{\partial t^*}{\partial \tau}=-\ln\left(1-\frac{V_{th}}{RI}\right)\) | positive |
| \(V_{th}\) | \(\displaystyle \frac{\partial t^*}{\partial V_{th}}=\frac{\tau}{RI-V_{th}}\) | positive |
| \(R\) | \(\displaystyle \frac{\partial t^*}{\partial R}=-\frac{\tau V_{th}}{R(RI-V_{th})}\) | negative |
Concrete point: with \(I=2,\tau=1,V_{th}=1,R=1\), \(t^*=0.6931\), while the derivatives are \(-0.5,\ 0.6931,\ 1,\ -1\) respectively.
具體數字:若 \(I=2,\tau=1,V_{th}=1,R=1\),則 \(t^*=0.6931\),四個 derivative 分別是 \(-0.5,\ 0.6931,\ 1,\ -1\)。
§5Factorial first-spike ordering spaceFactorial first-spike ordering space
The state is relational: no single neuron “is” the class. A state is the complete relative temporal order among all neurons.
state 是 relational:不是某一顆 neuron 代表 class,而是全部 neurons 的相對 first-spike 次序共同代表一個 state。
§7Navigator lossNavigator loss
The gradient is negative, so reducing loss means increasing every target-adjacent gap \(g_k\). If a pair is reversed, the loss produces a strong correction; if it is correct but too close to the boundary, it still receives a margin correction.
這個 gradient 永遠是負的,所以降低 loss 就是把 target-adjacent gap \(g_k\) 推大。pair 如果反了,修正很強;pair 雖然順序對但太靠近 boundary,仍會收到 margin correction。
A successful run need not end at zero loss. The formal 3.0 stopping rule stops at first strict-margin success, and softplus remains positive near the margin boundary.
成功不要求 final loss=0。3.0 formal run 在第一次達到 strict margin 時停止,而且 softplus 在 margin 附近仍為正。
§8Full gradient path: target → physical LIF parameter完整梯度鏈:target → physical LIF parameter
For a target-adjacent pair, \(g_k=t_b-t_a\). Therefore \(\partial g_k/\partial t_a=-1\) and \(\partial g_k/\partial t_b=+1\). A violated target relation directly pushes the earlier neuron earlier and/or the later neuron later.
對 target-adjacent pair,\(g_k=t_b-t_a\),所以 \(\partial g_k/\partial t_a=-1\)、\(\partial g_k/\partial t_b=+1\)。如果 target relation 錯了,gradient 會直接把該早的 neuron 推早、該晚的 neuron 推晚。
Physical values can be bounded by an unconstrained variable \(a_i\):
physical value 可用 unconstrained variable \(a_i\) 做 bounded map:
This is the learning mechanism: the Navigator does not classify 720 states; it creates a continuous temporal error whose gradient directly changes the LIF timing controls until the hard permutation flips into the target chamber. The optimizer only turns that gradient into numerical parameter updates; it is not the state-addressing representation.
這就是 learning mechanism:Navigator 不是做 720-class classification;它產生 continuous temporal error,gradient 直接改 LIF timing controls,直到 hard permutation 跨 boundary 進入 target chamber。optimizer 只負責把 gradient 變成數值更新,它不是 state-addressing representation 本身。
§9What is actually learnable?到底什麼是 learnable?
| arm | learnable values/run | meaning意義 |
|---|---|---|
| current | 6 × \(I_i\) | learnable external drivelearnable external drive;嚴格說不是 intrinsic neuron parameter |
| tau | 6 × \(\tau_i\) | intrinsic membrane time scaleintrinsic membrane time scale |
| threshold | 6 × \(V_{{th},i}\) | intrinsic firing thresholdintrinsic firing threshold |
| resistance | 6 × \(R_i\) | intrinsic current-to-voltage gainintrinsic current→voltage gain |
3.2 does not jointly train 24 values. It uses four separate matched arms. Each arm has only six trainable values. A joint \(I+\tau+V_{th}+R\) model would have 24 values, but that joint model has not been established.
3.2 不是同時學 24 個數字。它是四個獨立 matched arms;每個 arm 只有 6 個可學數字。若未來 joint \(I+\tau+V_{th}+R\) 才會是 24,但這個 joint model 尚未被建立。
“Learnable LIF parameter” means that the gradient directly updates a physical neuron parameter such as \(\tau_i,V_{th,i},R_i\). It does not yet mean that a reusable train-once addressing function has been learned.
「LIF parameter 可學」的精確意思是:gradient 直接更新 \(\tau_i,V_{th,i},R_i\) 這類 physical neuron parameter。它還不等於已經學會一個 train-once、可重複使用的 addressing function。
§10Experiment 3.0 — Direct Navigator addressingExperiment 3.0 — Navigator 直接 addressing
Question: can a frozen bounded-current LIF be pushed into an arbitrary target permutation using only the 5-D Navigator and gradient descent?
問題:固定 bounded-current LIF,只給 target permutation 與 5-D Navigator,能不能用 gradient descent 直接推進任意 target chamber?
| stage | runs | hard permutation | min g ≥ 0.1 | median steps | median loss |
|---|---|---|---|---|---|
| 3.0.0 pilot | 400 | 100% | 100% | 64 | 0.016274 |
| 3.0.1 exact coverage | 7,200 | 100% | 100% | 62 | 0.015158 |
3.0.1 covers all 720 target permutations with 10 initializations per target. All 7,200 runs pass both hard ordering and the strict \(\delta=0.1\) margin.
3.0.1 覆蓋全部 720 targets,每個 target 10 個 initializations;7,200/7,200 都通過 hard ordering 與 strict \(\delta=0.1\) margin。
Formal integrity: 5/5 equation/autodiff tests passed; independent recomputation of \((I,t,g,L)\) from final \(a\) matched to roughly \(10^{-15}\)-scale error. The formal mechanism used no NN/MLP, JEPA, oracle, double forward, factorial head, lookup, auxiliary loss, or weight decay.
Formal integrity:5/5 equation/autodiff tests 通過;從 final \(a\) 獨立重算 \((I,t,g,L)\) 的誤差約在 \(10^{-15}\) 等級。formal mechanism 沒有 NN/MLP、JEPA、oracle、double forward、factorial head、lookup、auxiliary loss 或 weight decay。
§11Experiment 3.1 — Direct state-to-state navigationExperiment 3.1 — 直接 state-to-state navigation
3.1 starts from a valid canonical source state and changes the target without resetting to a random initialization. All 518,400 ordered source→target transitions pass hard ordering and margin.
3.1 從合法 canonical source state 出發,不 reset 成 random initialization,直接換 target;518,400 個 ordered source→target transitions 全部通過 hard ordering 與 margin。
This is exact reachability without a supplied symbolic swap path, but not a shortest-path theorem. Only 291,039 / 517,680 nonidentity transitions (56.2199%) had sampled path length equal to Kendall shortest distance; 43.7801% had extra flips. Mean efficiency was 0.887137 and the minimum was 1/3.
這證明「不用 supplied symbolic swap path 也能 exact reach target」,但不是 shortest-path theorem。517,680 個 nonidentity transitions 中只有 291,039(56.2199%)的 sampled path length 等於 Kendall 最短距離;43.7801% 有額外 flips。mean efficiency=0.887137,最低=1/3。
The sampled Kendall distance was non-increasing at every sampled step in 0.586572 of transitions, so gradient navigation is exact but often non-monotone in hard permutation space.
sampled Kendall distance 每一步都不增加的比例只有 0.586572,所以 gradient navigation 最終 exact,但在 hard permutation space 裡常會繞路。
§12Experiment 3.2 — Intrinsic parameter controllabilityExperiment 3.2 — intrinsic parameter controllability
The Navigator and loss are unchanged. Only the physical parameter receiving the gradient is changed.
Navigator 與 loss 完全不變,只改「gradient 最後更新哪一個 physical parameter family」。
| arm | exact runs | hard + margin | median steps | median final loss |
|---|---|---|---|---|
| current reference | 3,600 | 1.0 / 1.0 | 62 | 0.01582082 |
| tau | 3,600 | 1.0 / 1.0 | 36 | 0.01799402 |
| threshold | 3,600 | 1.0 / 1.0 | 36 | 0.01627986 |
| resistance | 3,600 | 1.0 / 1.0 | 57 | 0.01791745 |
Each arm is 720 targets × 5 initializations = 3,600 runs. All four arms pass every target/initialization combination.
每個 arm 都是 720 targets × 5 initializations = 3,600 runs;四個 arms 全部通過。
| family | mean value by target rank 1→6target rank 1→6 mean | direction方向 |
|---|---|---|
| current | 4.451, 3.283, 2.528, 2.064, 1.756, 1.513 | decreasing |
| tau | 0.631, 0.932, 1.130, 1.323, 1.546, 1.919 | increasing |
| threshold | 0.479, 0.722, 0.877, 1.027, 1.185, 1.398 | increasing |
| resistance | 1.287, 1.103, 0.966, 0.856, 0.769, 0.688 | decreasing |
These directions exactly match the analytical derivatives: early ranks require larger \(I,R\) and smaller \(\tau,V_{th}\).
這些 rank-direction 和 exact derivative 完全一致:early rank 要 \(I,R\) 大、\(\tau,V_{th}\) 小。
§13Experiment 3.3 — Signed voltage-mediated couplingExperiment 3.3 — signed voltage-mediated coupling
3.3 freezes current and intrinsic parameters and learns only cross-node coupling:
3.3 固定 current 與 intrinsic parameters,只學 cross-node coupling:
Formal settings: \(dt=0.02\), horizon \(=2.5\), \(|W_{ij}|<0.08\).
formal 設定:\(dt=0.02\)、horizon \(=2.5\)、\(|W_{ij}|<0.08\)。
| arm | params/run | exact runs | hard + margin | median steps |
|---|---|---|---|---|
| row_shared | 6 | 2,160 | 1.0 / 1.0 | 42 |
| directed | 30 | 2,160 | 1.0 / 1.0 | 42 |
| target rank | row_shared incoming sum | directed incoming sum |
|---|---|---|
| 1 | +0.29943 | +0.30174 |
| 2 | +0.06805 | +0.07038 |
| 3 | −0.08717 | −0.08493 |
| 4 | −0.19592 | −0.19369 |
| 5 | −0.27491 | −0.27254 |
| 6 | −0.33481 | −0.33210 |
The evidence therefore supports rank-coded net incoming balance, not yet source-specific topology discovery. The 30-edge arm does not outperform the six-value row-shared arm in optimization speed. Across final directed edges, the reported sign fractions were about 0.32647 positive, 0.67352 negative, and 0.000015 near zero.
所以目前 evidence 支持的是 rank-coded net incoming balance,不是 source-specific topology discovery;30-edge arm 沒有比 6-value row-shared arm 更快。directed arm 的 final edge sign 比例約為 positive 0.32647、negative 0.67352、near-zero 0.000015。
Unlike 3.0–3.2, 3.3 does not have the same simple closed-form \(t^*(W)\). Gradients propagate through the numerical voltage-trajectory computational graph; this document does not invent an unsupported closed-form \(\partial t^*/\partial W\).
3.3 不像 3.0–3.2 有簡單 closed-form \(t^*(W)\)。gradient 沿 numerical voltage trajectory 的 computational graph 回傳;本文不虛構不存在的 \(\partial t^*/\partial W\) closed form。
§14What snnTorch snn.Leaky actually implementssnnTorch 的 snn.Leaky 到底在做什麼
The official snnTorch 1.0.0 documentation describes snn.Leaky as a discrete-time first-order LIF. With reset-by-subtraction:
官方 snnTorch 1.0.0 文件把 snn.Leaky 定義成 discrete-time first-order LIF。reset-by-subtraction 時:
Each call advances one timestep, so a normal forward simulation explicitly loops over time. The neuron returns spike and membrane state. The API supports learn_beta=True and learn_threshold=True.
每 call 一次只前進一個 timestep,所以一般 forward 會明確 loop over time;neuron 回傳 spike 與 membrane state。API 支援 learn_beta=True 與 learn_threshold=True。
Because the hard spike is a Heaviside step, typical snnTorch training uses a surrogate derivative in backward. The documentation uses ATan by default and shows alternatives such as fast sigmoid.
因為 hard spike 是 Heaviside step,典型 snnTorch training 會在 backward 用 surrogate derivative。官方文件預設使用 ATan,也示範 fast sigmoid 等替代方法。
Therefore “learnable decay” or “learnable threshold” alone is not the novelty of Gary's system; snnTorch already supports those ideas.
所以「decay 可學」或「threshold 可學」本身不是 Gary's system 的 novelty;snnTorch 已經支援。
§15Deep comparison: Gary's system vs snnTorch Leaky深度比較:Gary's system vs snnTorch Leaky
| aspect面向 | Gary's LIF Ordering System | snnTorch Leaky |
|---|---|---|
| time時間 | continuous exact first-spike in 3.0–3.2 | discrete timestep recurrence |
| primary output主要輸出 | \(t^*\rightarrow\operatorname{argsort}(t^*)\) | spk, mem per timestep |
| representationrepresentation | first-spike permutation state, nominal \(M!\) | spike trains / membrane dynamics; no built-in factorial state interpretation |
| gradient through spikespike gradient | 3.0–3.2 bypass binary spike and differentiate exact \(t^*\) | typically surrogate gradient for \(dS/dU\) |
| decay | continuous \(\tau_i\) | discrete \(\beta_i\), optionally learnable |
| threshold | \(V_{th,i}\), formally validated learnable arm | threshold, optionally learnable |
| membrane resistancemembrane resistance | explicit \(R_i\), formally validated arm | no explicit membrane-resistance API parameter; its effect is generally absorbed into input/current scaling沒有 explicit membrane-resistance API parameter;效果通常吸收到 input/current scaling |
| current | can itself be optimized as six values可直接作為 6-value optimization arm | input injection, often produced by previous weights/datainput injection,常由前層 weights/data 產生 |
| reset | irrelevant after the first spike for 3.0–3.2 ordering3.0–3.2 first spike 後 reset 不影響 ordering | subtract / zero / none are explicit recurrence options |
| target | \(\pi^*\rightarrow\) Navigator | user-defined task target/loss由使用者 task 定義 |
| loss | adjacent-margin softplus ordering loss | no built-in Navigator loss沒有內建 Navigator loss |
| typical use典型用途 | navigate a factorial relational state substrate導航 factorial relational state substrate | build SNN layers/networks for task learning建立 SNN layers/networks 做 task learning |
Core difference: snnTorch is a library for constructing and training spiking networks; Gary's current research asks whether first-spike ordering itself can be a factorial state substrate with a small differentiable addressing interface.
核心差異:snnTorch 是用來建立與訓練 spiking networks 的 library;Gary's system 現在研究的是 first-spike ordering 本身能不能成為 factorial state substrate,並由很小的 differentiable interface 去 address。
A conceptual relation between continuous time constant and discrete decay is \(\beta\approx e^{-\Delta t/\tau}\), but the exact physical mapping depends on discretization and input scaling, so they should not be treated as identical parameterizations.
continuous time constant 和 discrete decay 的概念關係常寫成 \(\beta\approx e^{-\Delta t/\tau}\),但 exact physical mapping 取決於 discretization 與 input scaling,所以不能把兩者當逐項完全一樣。
§16snnTorch RLeaky vs Experiment 3.3snnTorch RLeaky vs Experiment 3.3
The official snn.RLeaky recurrence adds a recurrent function of output spikes:
官方 snn.RLeaky recurrence 加的是 output spikes 的 recurrent function:
Experiment 3.3 instead couples subthreshold membrane voltages:
Experiment 3.3 則是 coupling subthreshold membrane voltages:
| difference差異 | 3.3 | RLeaky |
|---|---|---|
| coupling source | membrane voltage \(V_j\) | output spike \(S_{out}\) |
| interaction before first spike?first spike 前可互相影響? | yes | pure spike-feedback term has no spike feedback before anyone spikespure spike-feedback term 在所有人都還沒 spike 前沒有 spike feedback |
| can alter first winner?可改第一名? | yes, through pre-spike voltage coupling | not through the spike-feedback term before a first spike exists在第一個 spike 出現前,spike-feedback term 本身不能 |
| learnable recurrence | 6 row balances or 30 signed directed edges | all-to-all linear/conv recurrence or elementwise recurrent weight; learnable recurrent weights supported |
So 3.3 is not just “RLeaky with a new name.” It specifically studies voltage-mediated pre-spike coupling as a controller of first-spike ordering.
所以 3.3 不是「RLeaky 換名字」;它研究的是 voltage-mediated pre-spike coupling 如何控制 first-spike ordering。
§17Parameter counts and claim boundary參數量與 claim boundary
| mechanism | M6 learnable values/run | formal statusformal status |
|---|---|---|
| current-only | 6 | PASS |
| tau-only | 6 | PASS |
| Vth-only | 6 | PASS |
| R-only | 6 | PASS |
| joint I+tau+Vth+R | 24 | NOT ESTABLISHED |
| 3.3 row_shared | 6 | PASS |
| 3.3 directed | 30 | PASS |
Established已建立
- establishedM6 nominal factorial state substrate: 720 first-spike permutations.M6 nominal factorial state substrate:720 個 first-spike permutations。
- established5-D direct gradient addressing of all 720 targets.5-D Navigator 直接 gradient-address 全部 720 targets。
- established518,400 / 518,400 ordered source→target transitions, without supplied swap paths.518,400 / 518,400 ordered source→target transitions,不需 supplied swap path。
- establishedEach of I, tau, threshold, resistance independently serves as a six-value actuator.I、tau、threshold、resistance 各自都能獨立成為 6-value actuator。
- establishedVoltage-mediated signed coupling can also control the full M6 target set.voltage-mediated signed coupling 也能控制完整 M6 target set。
Not established尚未建立
- openTrain-once reusable addressing function for unseen targets.train once 後對 unseen target 立即 output parameter 的 reusable addressing function。
- openM>6 optimization scaling.M>6 optimization scaling。
- openRobustness under current/parameter/noise/spike-time perturbations.current / parameter / noise / spike-time perturbation 下的 robustness。
- openJoint-parameter identifiability or biological plasticity rule.joint-parameter identifiability 或 biological plasticity rule。
- openA functional task that actually uses all 720 states.真正 functional task 實際利用全部 720 states。
Recommended wording: “We establish an M6 factorial first-spike ordering substrate and an \(M-1\)-dimensional differentiable Navigator that, through per-instance gradient optimization, can control only \(M\) LIF timing values to address arbitrary target permutations.”
建議說法:「我們建立了一個 M6 factorial first-spike ordering substrate,以及一個 \(M-1\) 維 differentiable Navigator;透過 per-instance gradient optimization,只控制 \(M\) 個 LIF timing values 就能 address 任意 target permutation。」
§18Conclusion and references小結與 references
One-sentence mental model: Gary's system is not a new base neuron equation; it is a factorial first-spike state representation plus a five-dimensional temporal Navigator that directly differentiates into small LIF physical controls.
一句話心智模型:Gary's system 不是新的基本 neuron equation;它是「factorial first-spike state representation + 5-D temporal Navigator」,而 Navigator 的 gradient 直接進入少量 LIF physical controls。
| series | formal job | formal workloadformal workload | audit |
|---|---|---|---|
| 3.0 | lif30-adjmargin-formal-0827-r1 | 400 pilot + 7,200 exact target runs | PASS |
| 3.1 | lif31-navigation-formal-0828-r1 | 720 source builds + 100 pilot + 518,400 transitions | PASS |
| 3.2 | lif32-intrinsic-formal-0828-r1 | 800 pilot + 14,400 exact arm-runs | PASS |
| 3.3 | lif33-coupling-formal-0828-r1 | 200 pilot + 4,320 exact arm-runs | PASS |
The snnTorch comparison was checked against the current official documentation surfaced as snnTorch 1.0.0. Software APIs can change; the links above are the authority for future revisions.
snnTorch 比較已依目前官方文件(顯示為 snnTorch 1.0.0)核對。software API 之後可能變動,未來更新以官方 links 為準。