Neural Networks in Trading: From Transformers to Spiking Neurons (SpikingBrain)
Introduction
Financial markets have always been a place where chance coexists with regularity, and an immediate reaction is often more important than lengthy analysis. They have been compared to the ocean, where a calm surface can give way to a storm in an instant, and to an organism, where nerve impulses determine the rhythm of life. But perhaps the most accurate metaphor remains the nervous system, capable of responding to external stimuli with lightning speed and pinpoint precision. That is precisely why the emergence of the SpikingBrain framework, presented in the paper “SpikingBrain Technical Report: Spiking Brain-inspired Large Models,” has generated such keen interest.
It is inspired by brain biology and is based on the spiking model of information processing. Unlike traditional neural networks, which continuously churn through raw data, spiking networks act as sentinels. They remain silent until a critical moment arrives, and only then do they send a signal. This logic strikingly echoes the very nature of the market, where the price can fluctuate within a narrow range for hours, and then a single event upsets the balance — and the movement grows into a full-fledged trend.
To get a feel for this analogy, just take a look at Renko charts. Here, a new Renko brick appears not based on time, but only when the price reaches a certain threshold. All the minor noise disappears, leaving a clear pattern of movement. A spiking neuron works the same way. It remains silent until the signal being analyzed reaches a specified level. For a trader, this is akin to a clear market signal amid the chaos of quotes, where countless extraneous details obscure the essence.
In reality, the financial market is driven not by time, but by events. The example of the Non-Farm Payrolls report is particularly telling. A few minutes before the data is released, the market may seem sleepy. Quotes freeze, liquidity shrinks, and traders lie low in anticipation. But once the figures turn out to be unexpected, an explosion begins. Currency pairs make sharp jumps, spreads widen, and stop orders in the order book are triggered one after another. Conventional models often need time to reflect such abrupt changes. They average and smooth the data, and therefore detect the movement only when it is almost over. Spiking architecture captures such events differently: each impulse instantly reconfigures the network, allowing a trading decision to be made in the very same second while it still makes sense.
A similar pattern is observed during breakouts of psychological levels. The price hovers around the level for a long time; traders wait, and tension builds in the air. Then there is a single surge, and the level is broken. An avalanche of orders begins, the movement accelerates, and the market explodes. At such a moment, it is especially important that the model reacts specifically to the breakout event, rather than to the gradual accumulation of statistics. This behavior is entirely consistent with the intuition of an experienced trader. He knows that all the preceding noise was merely a prelude, and that the signal emerges precisely at the moment of the decisive move.
There is another example as well: the chain triggering of stop-losses. The price reaches a critical zone, and orders are activated one after another, creating an avalanche effect. This phenomenon is remarkably similar to the way a neural circuit works, where the first impulse triggers the second, the second triggers the third, and the wave continues to spread further and further. Traditional models tend to perceive such dynamics as chaotic movement. But for the spiking model, this is a natural process. Each spike becomes a trigger for a new one, and the system organically reflects such avalanche-like phenomena.
From these observations, it becomes clear that SpikingBrain thinks in terms of events. Unlike traditional models, which view the data flow as a continuous process, it highlights precisely those moments that truly carry meaningful information. Its operation is based not on constant analysis, but on immediate reaction. This approach avoids wasting resources on noisy fluctuations and focuses on what truly drives the market.
Energy efficiency is another key advantage. In traditional models, neurons are always active, even when there is no meaningful input signal. In spiking models, the situation is different. A neuron fires only when the threshold is reached. In market conditions, this means that computing resources are used only at the moment of an important movement. This approach makes it possible to run the algorithm directly in trading terminals, without connecting to external servers and without excessive delays.
Another quality of SpikingBrain that makes it particularly valuable in practice is its capacity for online learning. The market is constantly changing, and patterns that worked yesterday may prove useless tomorrow. Classical models often find themselves held hostage to the past. They are trained on old data and struggle to handle unexpected situations. The spiking model, by contrast, adjusts its forecast on the fly. Each new event becomes part of its learning process, allowing the algorithm to remain flexible and dynamic. This closely resembles the behavior of an experienced trader who reacts instantly to the latest news and adjusts their strategy without getting stuck in old patterns.
All these qualities make it possible to view SpikingBrain as an attempt to build a full-fledged nervous system for a trading robot. It thinks in impulses, reacts to events, and adapts to new conditions just as quickly as the market does. Much like Renko charts, which help traders see the essence of price movements without unnecessary noise, the spiking model highlights what matters and ignores what is secondary.
The SpikingBrain Algorithm
If we look at SpikingBrain through the lens of the evolution of neural network concepts, it appears to be a natural step forward — both logical and timely. The Transformer once sparked a true revolution by moving away from cumbersome recurrent structures and introducing an attention mechanism capable of keeping the entire data sequence in focus at once. This versatility has made it the gold standard for language models and a wide range of related tasks. However, versatility also has its downside. An architecture that works perfectly with text quickly becomes overloaded when it encounters the market. After all, most of the time nothing significant happens, and there are too many noisy fluctuations.
This is exactly where SpikingBrain brings a fresh idea to the table. The framework's authors draw on the concept of attention, but shift it into the event-driven domain. In a classic Transformer, attention is distributed across the entire sequence — the model considers everything at once, identifying parts that are more or less significant. In SpikingBrain, attention arises like a flash — an instantaneous response when a threshold level is reached.
This difference is particularly noticeable in the financial market. Quotes can be thought of as a stream of letters that form the words and sentences of the market's history. Unlike natural language, there are a huge number of extra letters, repetitions, and random insertions here. The Transformer is capable of processing them all. But it often does so excessively, like a pedantic teacher who does not overlook a single detail, even a meaningless one. SpikingBrain, on the other hand, acts like an experienced editor, cutting out predictable patterns and focusing only on what changes the meaning. The result is that the data structure remains streamlined, and the signal is not overloaded with unnecessary information.
In addition, energy efficiency is important. The Transformer requires enormous computational resources. This is justified for tasks such as text generation or translation, where every detail matters. In the financial market, this kind of overanalysis leads to false signals and unnecessary costs. The spiking model works differently. A neuron remains silent until it has accumulated enough potential, and only then does it send a signal. That is how the model conserves resources. Its behavior becomes closer to the nature of trading. A trader does not react to every tick; instead, they wait for the moment when a price movement is truly worth paying attention to.
SpikingBrain can be viewed as a hybrid. On the one hand, it inherits the conceptual power of the Transformer; on the other, it enriches it with biologically inspired event-driven dynamics. In this reinterpretation, attention is no longer a background element but an act of choice — an instantaneous shift of the system into a reactive mode. It is precisely this property that makes the model particularly valuable for a market where the decisive role is played not by all data, but only by critical moments capable of altering the trajectory of capital flows.
Saving computational resources is a compelling argument in and of itself. However, in financial markets, it also has a deeper meaning. It is not just about speed, but also about the ability to produce timely, reliable signals. If a model is overloaded with excessive information, it either lags or generates false signals. The quadratic complexity of the Transformer, as the sequence length increases, causes the system to start choking. Instead of a clear picture, the trader gets noise disguised as analysis.
SpikingBrain tackles this problem in a radically different way. Event-driven processing reduces the load and naturally filters the data. A neuron fires only when the accumulated dynamics actually matter. The model turns into a kind of market editor, weeding out meaningless fluctuations and retaining only the turning points. As a result, the signals become less frequent but more reliable — exactly what is needed in trading, where every decision involves risk and must be well-founded.
In addition, reducing the computational load opens the door to more complex application scenarios. Where the classic Transformer runs into hardware limitations, the spiking architecture can process large volumes of market data in real time without sacrificing performance. For algorithmic trading, where milliseconds can determine the outcome of a trade, this advantage is critical.
The SpikingBrain architecture is built on principles that organically combine spiking processing with an attention mechanism reworked for event-driven operation. It is based on a set of spiking neurons, each of which receives an input signal and accumulates potential until it reaches a critical threshold. At that moment, the neuron fires, generating a spike that is transmitted further through the network, affecting other neurons and activating the corresponding processing units.
This approach allows the network to focus on significant events while ignoring minor fluctuations and noise-driven variations. The attention mechanism in SpikingBrain is no longer a distributed weight, as in the classic Transformer, but instead becomes a dynamic filter. It amplifies the response of neurons that receive the most relevant input and attenuates the response to less significant signals. As a result, the network does not react to every change, but only to those moments that are actually capable of altering the market situation.
Connections between neurons are structured to facilitate both local and global responses. Local connections track changes in individual data segments. Global connections allow signals to travel farther, triggering synchronization and responses to larger-scale market events, such as a breakout of a liquidity level or the avalanche-like triggering of stop orders. This structure makes it possible to retain detailed information while maintaining control over the overall picture of the market.
Another key element is threshold adaptation. The threshold for neuron activation is not fixed but changes dynamically depending on current volatility and historical dynamics. This allows the model to remain sensitive to unexpected movements while preventing an excessive response to minor noise. Combined with event-driven processing, this makes the SpikingBrain architecture particularly resilient to market fluctuations and capable of generating high-value signals.
In their paper, the authors of the SpikingBrain framework conclude that, through lightweight fine-tuning, the basic Transformer can be transformed into various effective attention variants. This opens up opportunities for flexible trade-offs between accuracy and computational efficiency. This is particularly important in the context of financial markets: some tasks require rapid analysis of long price sequences, while others require high signal accuracy with limited resources.
To illustrate this, the authors’ paper examines two approaches implemented using a single pre-trained classical Transformer model. The first, SpikingBrain-7B, is a purely linear model optimized for processing long contexts with minimal resource consumption. The second, SpikingBrain-76B, a hybrid MoE variant, combines linear and local attention, striking a balance between efficiency and prediction quality.
In SpikingBrain-7B, linear attention alternates with local layers of sliding-window attention with a fixed analysis window. Local layers capture detailed price patterns, much like a trader identifies short-term impulses, whereas linear attention efficiently condenses long-range dependencies, allowing the model to preserve contextual integrity without excessive computational overhead. As a result, model training remains linear in time, while memory usage during operation remains constant regardless of the sequence length. In practical terms, linear attention is implemented through the Gated Linear Attention module, which amplifies important signals and helps neurons remember market dynamics by responding only to truly significant events.
SpikingBrain-76B takes this idea a step further. In this architecture, linear and local attention are combined within a layer in a parallel hybrid configuration. Standard full-attention layers are also inserted between these layers to enhance the global coherence of the signals. This architecture makes it possible to capture both local and global market events simultaneously without overloading the model with unnecessary data. To increase the flexibility of local attention and reduce the attention-sink effect, special trainable tokens have been added to the model; these tokens interact with the sequence without causal constraints. In addition, a sparse Mixture-of-Experts architecture is used here for the FFN modules. Each layer activates only a subset of the experts, which conserves resources and enables the model to process large volumes of data in real time.
Both implementations demonstrate how Transformer principles can be reinterpreted within an event-driven paradigm.
The SpikingBrain architectural design is, in many ways, aligned with the operating principles of the biological brain. Linear attention modules exhibit properties similar to those of human memory. They rely on compressed and continuously updated states. At each step, they extract information only from the current memory state, exhibiting behavior reminiscent of a Markov process. From a biological perspective, this type of recurrent processing over time can be viewed as a simplified abstraction of dendritic dynamics with a multi-branched morphology, in which each branch performs local signal processing.
The Mixture-of-Experts (MoE) component embodies the principle of modular sparse activation and functional specialization, which is similar to distributed and specialized processing in neural circuits. Each expert processes only a portion of the information, and the overall result is assembled from active modules, which enhances the system's efficiency and stability.
The spike encoding scheme proposed by the framework’s authors is inspired by biological systems characterized by event-driven and adaptively sparse neuronal activation. By combining model-level sparsity (MoE) with sparsity at the individual-neuron level, computational resources are allocated efficiently, creating a two-level efficiency mechanism. This approach allows the system to activate only as needed, focusing resources on the most significant market events while maintaining high data processing speeds.
Taken together, these solutions demonstrate a promising path toward creating architectures for large models that are both efficient and computationally economical. At the same time, they maintain biological plausibility by combining the power of neural networks with the organizational principles of the living brain. In the context of financial markets, this means that SpikingBrain is capable of highlighting key events while acting selectively and minimizing noise.
Training SpikingBrain is based on principles that enable effective knowledge transfer from pre-trained Transformer models to new spiking and hybrid architectures. This is based on the correspondence between attention maps. In a pre-trained Transformer, attention is determined by the standard SoftMax function applied to the query-key product multiplied by the sequence mask. In this context, local sliding-window attention (SWA) can be viewed as a sparse version of this map with a strong bias toward recent events. In that case, linear attention is a low-rank approximation, where the maximum rank of the map is limited by the dimensionality of the query and key vectors.
Using this correspondence, the QKV projection parameters can be initialized directly from a checkpoint of a pre-trained Transformer. Through lightweight fine-tuning on a relatively small amount of data, attention is adapted to local or low-rank settings. Local attention preserves the accuracy of the analysis of recent events, while linear attention captures global relationships. Their combination in hybrid mode provides a more accurate approximation of the original attention map. For the financial market, this means that the model can simultaneously track short-term price spikes and global trends while maintaining consistency in its signals.
Several important techniques are used to ensure stable training and a smooth transition to the new architecture. First, QK activations in linear attention are converted to non-negative values using ReLU or Sigmoid to preserve the properties of SoftMax and ensure correct knowledge transfer. Second, low-rank parameters, such as normalization and gating components, are retained. During the conversion stage, training is performed with a small step size, which makes it difficult to optimize a large number of randomly initialized parameters. The framework authors reuse all projection weights in the attention modules and FFN, thereby minimizing the number of new parameters.
The third feature is the extension of the long context during fine-tuning. Since efficient attention mechanisms scale subquadratically, it is practical to limit the context length during the initial conversion and gradually increase it during training. At the same time, computational efficiency is maintained. The model is fully trained during the conversion phase to ensure performance, using learning rate warm-up and distillation techniques between architectures, or by training all parameters at once.
In a financial context, this approach enables the smooth transfer of knowledge from powerful pretrained models to specialized spiking SpikingBrain architectures. The model retains information about key market events while effectively filtering out noise and reducing computational costs.
The SpikingBrain training process is implemented as a multi-stage conversion. Continual Pre-Training (CPT), with gradual context expansion, smoothly transitions into Supervised Fine-Tuning, SFT. In their work, the authors of the model training framework used high-quality open datasets. Nevertheless, in real-world practice, specific financial data may be used to adapt the model to particular markets or assets.
The authors’ article describes a three-stage process for training the SpikingBrain language model. At each of these stages, the context length gradually increases, and the model adapts to the new requirements. In the first stage, the models are trained on 100 billion tokens with a sequence length of 8K. This makes it possible to transfer attention patterns to the local and low-rank variants and achieve stable convergence of the loss function. In the second and third stages, the sequence length is extended to 32K and 128K, respectively, using 20–30 billion tokens each time. In total, the conversion requires about 150 billion tokens, which is only about 2% of the data needed to train the model from scratch, and provides effective adaptation with limited resources. At all stages, the base positional RoPE encoding remains unchanged, as in the original model.
The CPT stage is followed by fine-tuning, which is also divided into three stages. The first of these focuses on basic language understanding and domain knowledge, using the extensive Infinity Instruct dataset, which covers scientific knowledge, code interpretation, and mathematical problem solving. Training is conducted on 500,000 examples with a sequence length of 8K. The second stage focuses on dialogue skills and instruction following, using a set of task-oriented multi-turn dialogues and questions with educational content. The data volume and sequence length remain the same. The third stage focuses on reasoning tasks, using a high-quality dataset containing detailed reasoning chains for mathematical proofs, logical inferences, case analysis, and other multi-step problems.
This multi-stage approach allows SpikingBrain to gradually acquire knowledge from pre-trained models, adapt to long contexts, and simultaneously improve signal accuracy at various levels.
To effectively expand the dense FFN model into an MoE architecture, the authors of the framework employ the upcycling technique, which increases the model’s capacity while reusing the knowledge already encoded in the original parameters. During initialization, the parameters of the base dense FFN model are copied to N experts, and a random router is introduced for each token. This router selects the Top-K experts with probability p and returns their weighted sum, ensuring functional equivalence to the original dense network at initialization.
During training, stochastic routing and the addition of noise to the data gradually break the initial symmetry, creating differentiated gradients and promoting expert specialization. In addition to the routed experts, the SpikingBrain-76B model uses a single shared expert that is always active for all tokens and helps stabilize the conversion process.
Direct copying and activation of multiple experts increases the scale of the output values. To maintain consistency between the outputs of MoE and the dense FFN, a scaling factor is introduced and applied to all experts during initialization. Thus, initially, the MoE network operates equivalently to the original dense model, while retaining the potential for further specialization and increased capacity.
In a financial context, this MoE architecture allows the SpikingBrain model to simultaneously monitor many market scenarios, assigning specialized experts to different types of signals. Some focus on short-term price impulses, while others focus on global trends or rare events. Thanks to upcycling, the network adapts quickly by reusing previously accumulated knowledge, which speeds up training and improves signal accuracy without imposing excessive computational load.
Inspired by biological computational mechanisms — event-driven processing and sparse activation — the authors of the framework propose a specialized spiking encoding strategy that converts the activations of large models into integer values and sequences of spikes. This scheme can be applied both during and after training, converting the model’s activations into spiking signals. To improve energy efficiency, alongside spiking, it is proposed to quantize the model weights and the KV cache to INT8 precision. Integrating this strategy with the lightweight SpikingBrain conversion eliminates the need for full fine-tuning. A relatively small calibration sample is sufficient to optimize the quantization parameters.
Activations in SpikingBrain are converted according to a two-stage logic. In the first stage, adaptive threshold spiking is used. Activations are converted into integer spikes using a dynamically regulated threshold that maintains a statistical balance of neural activity. This prevents both excessive neuron firing and inactivity, which is critical for preserving the information content of the signals.
Formally, the threshold is defined as the average absolute value of the membrane potential, scaled by the hyperparameter k, which controls the number of spikes. This mechanism is resistant to rare outliers. Neurons maintain stable activity for most signals and increase the number of spikes only for critically significant surges. In a financial context, this can be compared to a price crossing a moving average line. Ordinary price fluctuations do not reach the neutral threshold and therefore trigger no response. Only when the price crosses the moving average — the key moment — is a signal generated. Similarly, a spiking neural network responds only to significant events, filtering out noise and focusing on market movements that have real trading value.
In the second stage, already at the inference stage, the resulting integer spikes are unfolded along the time axis into sparse spike sequences with values {−1, 0, 1}. This makes it possible to replace dense matrix multiplications with event-based accumulations, significantly improving computational efficiency. The framework authors propose three encoding schemes for different scenarios.
The first is binary encoding {0,1}. Each one corresponds to a neuron activation, and the number of spikes is accumulated over time. This is a simple and effective solution for low spiking density; however, large values require many time steps.
The second is ternary encoding {−1, 0, 1}, which introduces inhibitory spikes. This approach allows both excitation and inhibition to be expressed, which is closer to the biological principles of excitation/inhibition. Ternary encoding reduces the number of time steps and cuts the spiking rate by more than half while preserving the expressiveness of the signals.
The third is bitwise encoding (bitwise coding). The integer number of spikes is unfolded bitwise over time. This scheme ensures maximum compression along the time axis. Bidirectional encoding and signed representation are possible, allowing both positive and negative values to be accounted for simultaneously while maintaining simplicity and biological plausibility. Bitwise encoding reduces the total volume of spikes by up to 8 times, minimizing communication overhead and computational load.
As a result, this approach allows SpikingBrain to avoid reacting to every tick or noisy market movement and instead focus on significant events — much like a trader who makes a decision only when the price crosses a key level. Each conversion of activations into spikes makes the model more selective, efficient, and better adapted to the dynamic environment of financial markets.
The authors’ visualization of the SpikingBrain framework is shown below.

Implementation in MQL5
After a detailed examination of the theoretical aspects of the SpikingBrain framework, we move on to the practical part of our work, where we will explore one possible implementation of the proposed approaches using MQL5. And we should take off our rose-colored glasses right away. The resource savings mentioned by the framework's authors are largely achieved through the use of pre-trained large language models based on the Transformer architecture. Without access to such models, the training process remains labor-intensive and resource-intensive, and it is important to take this into account when planning your work.
Nevertheless, the key question regarding the effectiveness of the proposed architecture remains open. This is precisely what we intend to evaluate in practice by examining how spiking processing and the modular structure of MoE perform under conditions of limited computational resources and real-world market data. This will allow us to assess the real advantages of SpikingBrain: how capable it is of filtering out noise, identifying key market signals, and conserving computational resources in an environment where milliseconds and precision are critical.
At the same time, the use of pre-trained neural models, as proposed by the framework’s authors, opens up an important opportunity. We can utilize previously built modules and layers of the neural network in this implementation. This significantly speeds up the process. However, it will not be possible to do entirely without implementing new algorithms. The combination of off-the-shelf components and new algorithms will make it possible to effectively implement SpikingBrain using MQL5 and test its effectiveness.
Let's start by developing a relatively simple algorithm that converts the continuous signal of a standard neuron into a discrete ternary value {-1, 0, 1}. This approach follows the basic idea of spiking activation: a neuron remains silent until the potential reaches a critical threshold, and only then does it fire a signal in either the positive or negative direction. To implement this algorithm, we create a new kernel in the OpenCL program.
The way the kernel works is simple and straightforward. First, we check that the source value is valid and is not NaN or infinity. If there is no signal or if the signal is too weak relative to the threshold level, the neuron remains in a resting state — it is assigned a value of 0.
__kernel void FloatToSpike(__global const float* values, __global const float* levels, __global float* outputs ) { const size_t id = get_global_id(0); float val = IsNaNOrInf(values[id], 0.0f); if(val == 0.0f) outputs[id] = 0.0f; else { const float lev = IsNaNOrInf(levels[id], 0.0f); if(fabs(val) < lev) outputs[id] = 0.0f; else outputs[id] = (float)sign(val); } }
If the absolute value exceeds the threshold, the neuron fires a spike based on the direction of the change: +1 for a positive impulse and −1 for a negative one.
This approach makes it possible to model the fundamental principle of spiking processing — responding only to significant events. Converting the data to {-1, 0, 1} makes it possible to filter out insignificant movements and focus on key event points, just as a trader reacts only when the price crosses an important level.
However, behind the apparent simplicity of the forward pass algorithm lies the nontrivial task of distributing the error gradient required to train the model. The main question is when and how the gradient should be propagated back to the neuron's input signal. If we limit gradient propagation to active neurons only, we can easily run into the problem of mass deactivation of a large number of neurons, at which point training effectively comes to a halt. The gradient simply does not propagate, and the model stops adapting.
The complexity is further compounded when using a three-valued signal {-1, 0, 1}. During the learning process, it inevitably becomes necessary to change the direction of a neuron’s activation from positive to negative — or vice versa — by passing through a range of inactivity. If the gradient is blocked within this range, the neuron freezes and loses its ability to learn.
Based on this, we make a fundamental decision: the error gradient must be propagated at all times, regardless of the neuron's current state. This approach ensures continuous parameter updates, prevents training from stalling, and preserves the model's adaptability, even if most neurons remain temporarily inactive. This is a key step that ensures the effectiveness of spiking processing when working with discrete signals and enables the model to learn stably from financial data streams.
The second important aspect concerns the threshold at which neurons fire. As the authors of the SpikingBrain framework note, a high threshold results in infrequent signal generation. As a result, a significant portion of useful information passes unnoticed by the model. Conversely, a low threshold increases neuron sensitivity, but at the same time makes them respond to noise, generating false signals.
Consequently, the optimal solution is to allow the model to learn the firing thresholds for each neuron on its own. This gives rise to an interesting property: the error gradients for the signal and the firing threshold are directed in opposite directions. If it is necessary to amplify a neuron's response to significant events, the threshold should be lowered so that the signal can more easily cross the barrier. If, on the other hand, the goal is to suppress noisy fluctuations, the threshold must be raised to limit false signals.
Thus, training threshold values becomes a delicate balancing act. The model must simultaneously filter out noise and remain sensitive to key market events. This mechanism makes spiking neurons adaptive and enables them to focus on truly significant price fluctuations. Like an experienced trader who knows when to act and when to wait for confirmation.
We move the implementation of the described approach to the kernel level of the OpenCL program, creating an algorithm that processes the error gradients of neuron signals and their threshold values.
__kernel void FloatToSpikeGrad(__global const float* values, __global float* values_gr, __global float* levels_gr, __global const float* gradients ) { const size_t id = get_global_id(0); const float grad = IsNaNOrInf(gradients[id], 0.0f); values_gr[id] = grad; if(fabs(grad) > 0.0f) { float val = IsNaNOrInf(values[id], 0.0f); levels_gr[id] = (float)(-sign(val) * grad); } else levels_gr[id] = 0.0f; }
The kernel's operating principle reflects a key feature of threshold neuron training. The error gradient is always propagated to the neuron's input signal via the values_gr array. In this case, for the neuron's threshold value, the gradient is formed with the opposite sign relative to the direction of the signal. If it is necessary to enhance a neuron's response to a significant event, the threshold is lowered. And when suppressing noise, it is raised.
To organize the process on the main program side, we create a new CNeuronSpikeConv layer, which inherits the basic functionality of fully connected objects while allowing the classical convolution process to be combined with the generation of spikes, forming a single module that produces discrete pulses for further processing.
class CNeuronSpikeConv : public CNeuronBaseOCL { protected: CNeuronConvOCL cConv; CParams cLevels; //--- virtual bool FloatToSpike(void); virtual bool FloatToSpikeGrad(void); //--- virtual bool feedForward(CNeuronBaseOCL *NeuronOCL) override; virtual bool updateInputWeights(CNeuronBaseOCL *NeuronOCL) override; virtual bool calcInputGradients(CNeuronBaseOCL *NeuronOCL) override; public: CNeuronSpikeConv(void) {}; ~CNeuronSpikeConv(void) {}; //--- virtual bool Init(uint numOutputs, uint myIndex, COpenCLMy *open_cl, uint window, uint step, uint window_out, uint units_count, uint variables, ENUM_OPTIMIZATION optimization_type, uint batch); //--- virtual CBufferFloat* GetWeightsConv(void) { return cConv.GetWeightsConv(); } virtual int Type(void) override const { return defNeuronSpikeConv; } //--- methods for working with files virtual bool Save(int const file_handle) override; virtual bool Load(int const file_handle) override; virtual void SetOpenCL(COpenCLMy *obj) override; //--- virtual bool WeightsUpdate(CNeuronBaseOCL *source, float tau) override; virtual uint GetFilters(void) const { return cConv.GetFilters(); } virtual uint GetVariables(void) const { return cConv.GetVariables(); } virtual uint GetUnits(void) const { return cConv.GetUnits(); } };
The class contains several key components. Specifically, the cConv object performs signal convolution operations, while cLevels learns and stores threshold values for each neuron. The class defines the FloatToSpike and FloatToSpikeGrad methods, which are wrapper methods for the kernels of the same name.
A new object is initialized in the Init method, which sequentially configures all of the layer's key components.
bool CNeuronSpikeConv::Init(uint numOutputs, uint myIndex, COpenCLMy *open_cl, uint window, uint step, uint window_out, uint units_count, uint variables, ENUM_OPTIMIZATION optimization_type, uint batch) { if(!CNeuronBaseOCL::Init(numOutputs, myIndex, open_cl, window_out * units_count * variables, optimization_type, batch)) return false;
First, the parent object is initialized, during which the inherited base interfaces are created.
Next, the convolution module cConv is initialized, and the parameters for the window, stride, number of elements, and variables are passed to it. The TANH activation function is set for it right away, which ensures smooth scaling of the output signal within the expected range of values.
if(!cConv.Init(0, 0, OpenCL, window, step, window_out, units_count, variables, optimization, iBatch)) return false; cConv.SetActivationFunction(TANH);
Next, the cLevels threshold component is configured, also with optimization enabled and the SIGMOID activation function defined. An important step is filling the weights of cLevels with constant values of -10. This ensures that the thresholds are set to a correct initial state before training — close to zero — allowing the neurons to respond immediately to significant events without being blocked by a high threshold.
if(!cLevels.Init(0, 1, OpenCL, Neurons(), optimization, iBatch)) return false; cLevels.SetActivationFunction(SIGMOID); CBufferFloat* temp = cLevels.getWeights(); if(!temp) return false; if(!temp.Fill(-10)) return false; //--- return true; }
As a result, the Init method brings together the preparation of all layer components, forming a ready-to-use module capable of both performing convolution and generating spiking pulses. This integration ensures that the layer processes financial data in a coordinated manner, allowing the neurons to respond only to significant fluctuations and filter out noise.
The feedForward method provides forward propagation of the signal through the CNeuronSpikeConv layer. First, the method of the same name in the cConv convolution module is called; it processes the input data and generates the convolved signal.
bool CNeuronSpikeConv::feedForward(CNeuronBaseOCL *NeuronOCL) { if(!cConv.FeedForward(NeuronOCL)) return false; if(bTrain) if(!cLevels.FeedForward()) return false; //--- return FloatToSpike(); }
If the layer is operating in training mode (bTrain==true), an additional forward pass is performed by the threshold generation component cLevels, ensuring correct threshold adaptation for each neuron.
The final step is to call the FloatToSpike function, which converts the obtained values into discrete spiking pulses.
As you can see, the algorithm for the forward pass is fairly straightforward and does not create complications for gradient propagation. Model parameters are updated using nested components, which helps maintain architectural cleanliness and simplifies the training process. Therefore, we suggest leaving the backward-pass methods for independent study. The complete code for the class and all of its methods is provided in the attachment, allowing you to review the implementation of all details and the integration of spiking neurons.
We've done a lot of work today, and it's time to take a break to let our thoughts settle. Let the accumulated information fall into place, and give the details time to sink in. In the next article, we will resume the work we started, continue our exploration of spiking models, and demonstrate how these algorithms are applied to real financial data streams, transforming theory into practical trading signals.
Conclusion
In this article, we examined the theoretical foundations of the SpikingBrain framework. The main advantage of the framework lies in its event-driven nature. Neurons are activated only when a threshold level is reached, which drastically reduces the likelihood of false signals and improves signal quality. Unlike traditional architectures, such as Transformer, the framework enables critical information about significant market movements to be preserved without overloading the system with redundant data.
In addition, the use of adaptive thresholds and the ability to integrate pre-trained modules ensure the model’s flexibility and scalability. The spiking architecture naturally supports energy and computational efficiency, making it particularly attractive for algorithmic trading and real-time analysis of large volumes of data.
Links
Programs used in this article
| # | Name | Type | Description |
|---|---|---|---|
| 1 | Study.mq5 | Expert Advisor | Expert Advisor for offline model training |
| 2 | StudyOnline.mq5 | Expert Advisor | Expert Advisor for online model training |
| 3 | Test.mq5 | Expert Advisor | Model testing Expert Advisor |
| 4 | Trajectory.mqh | Class Library | Structure for describing the system state and model architecture |
| 5 | NeuroNet.mqh | Class Library | Class library for building a neural network |
| 6 | NeuroNet.cl | Library | Library of OpenCL program code |
Translated from Russian by MetaQuotes Ltd.
Original article: https://www.mql5.com/ru/articles/19709
Warning: All rights to these materials are reserved by MetaQuotes Ltd. Copying or reprinting of these materials in whole or in part is prohibited.
This article was written by a user of the site and reflects their personal views. MetaQuotes Ltd is not responsible for the accuracy of the information presented, nor for any consequences resulting from the use of the solutions, strategies or recommendations described.
Uncertainty as a Model (Part 3): Mathematical Statistics — How to Extract Knowledge from Data
Time Series Shapelets: Learning a Price Shape, and Testing Whether It Means Anything
Regime Discovery by Structure: Implementing Toeplitz Inverse Covariance Clustering (TICC)
Inside MetaEditor's AI Assistant: Writing, Repairing and Testing MQL5 with an Agent
- Free trading apps
- Over 8,000 signals for copying
- Economic news for exploring financial markets
You agree to website policy and terms of use