Neural Networks in Trading: From Transformers to Spiking Neurons (Key Components)
Introduction
Financial markets are like a stormy sea, where a ship’s captain must make split-second decisions, relying not only on charts and compasses but also on intuition honed through years of experience. In today's world, computational models are increasingly taking on this role, and the researcher's task is to create an algorithm capable of rivaling human flexibility and foresight. Against the backdrop of traditional machine learning methods and well-established neural network architectures, the SpikingBrain framework emerged as a breath of fresh air, capable of filling the researcher’s sails and leading them to new shores of data analysis.
A key advantage of SpikingBrain is its event-driven approach to information processing. Unlike classical models, in which data is treated as a continuous stream, spiking networks respond to discrete events, which “makes them better aligned with the nature of real financial signals. The exchange-traded markets never evolves in a steady manner: some periods are calm, with barely noticeable fluctuations, while others are like a squall, bringing sharp spikes in prices. A traditional neural network is forced to process both modes with equal intensity, overloading computational resources and losing accuracy during periods of sudden turbulence. SpikingBrain, on the other hand, allows the system to focus on events, giving it natural efficiency and adaptability.
Equally important is the fact that spiking models are energy-efficient. For financial markets, where algorithms often need to run continuously, the ability to reduce the load on hardware without compromising the quality of analysis is a compelling argument. In an era when computing resources are distributed across thousands of strategies and dozens of servers, the ability to maximize benefits while minimizing energy consumption takes on not only technical but also strategic significance. It could be said that SpikingBrain paves the way for intelligent computing. The focus is not on the raw power of the machines, but on elegant architecture and process optimization.
The framework's particular strength lies in its ability to support effective knowledge transfer and adaptation. Unlike bulky models that require massive training datasets to function reliably, the spiking architecture makes better use of previously accumulated experience and adapts more quickly to new data. This makes it possible to apply pre-trained models to financial analysis tasks, fine-tuning them even on relatively small datasets. This approach significantly speeds up the model’s practical deployment and makes it a flexible tool capable of accounting for the specific characteristics of a particular market. Consequently, SpikingBrain serves as a bridge between the versatility of pre-trained solutions and the unique characteristics of local data.
It is worth emphasizing that SpikingBrain is not merely a declaration of principles, but a coherent system grounded in careful engineering design. It is based on the idea of converting the input signals into a sequence of spikes, which serve as the building blocks for all subsequent processing. This makes it possible to form a rich representation of time series, where each burst of activity carries significant information about the state of the system.
Specialized blocks that mimic biological mechanisms are then responsible for transmitting and integrating these signals. They provide flexibility in working with various types of data and lay the foundation for building a multi-level analytical structure.
Here, it is worth taking a closer look at an important element of the architecture — the mechanism for encoding signals into spikes. For financial time series, this task is particularly challenging: it is not enough simply to record price or volume values; they must also be transformed into a format in which each spike reflects a significant change in the data. The authors propose several encoding strategies. From a simple threshold-based approach, where a spike occurs when a critical value is exceeded, to more complex schemes that take into account the rate of change and the context of previous events. This wide variety of methods makes the model versatile. It can perform equally well with high-frequency data, where every millisecond counts, and with quieter daily charts, where the overall trend is what matters most.
The next key block is the integrators. They are responsible for accumulating and summing incoming impulses, determining the moment when a neuron becomes active. From the perspective of financial markets, this is similar to a noise-filtering mechanism: many small fluctuations can be ignored, but their cumulative effect, once it reaches a certain threshold, triggers a response from the model. This principle allows the model to maintain a balance between sensitivity and stability, which is particularly valuable in the face of market noise.
The framework’s authors proposed a model in which the hybrid organization of layers plays a central role. Some of them specialize in identifying short-term patterns, capturing rapid market fluctuations. Others focus on integrating signals over a longer time horizon, which makes it possible to generate more stable forecasts. This combination gives the system a kind of dual vision. It is capable of both capturing the finest details of the present moment and maintaining a strategic perspective. For financial markets, this is akin to being able to see both the movement of a stopwatch’s second hand and the overall rhythm of the pendulum that determines the passage of time.
Modules responsible for normalization and network stability are an important part of the architecture. Unlike traditional neural networks, where balance is achieved through extensive regularization procedures, SpikingBrain uses more subtle mechanisms embedded in the very nature of spiking signals. As a result, the network remains stable even when faced with chaotic data, while retaining its ability to make accurate predictions. For the financial market, this quality cannot be overstated: after all, chaos is its constant companion, and the ability to identify hidden patterns within it determines the success of any model.
It is also worth noting the modularity of the proposed architecture. The authors have laid down principles that make it possible to expand the model both in breadth and depth, combine spiking elements with traditional deep learning blocks, and develop hybrid solutions. This synthesis opens up even more possibilities for applications in finance. For example, it is possible to build models in which spiking mechanisms are responsible for processing event-driven signals, while classical blocks handle long-term forecasts or the analysis of textual information. The result is a multifaceted tool in which each part of the system performs its own function, and together they form a comprehensive mechanism for analyzing a complex market landscape.
The authors’ visualization of the SpikingBrain framework is shown below.

In the practical section of the previous article, we focused on laying the groundwork necessary for transferring the ideas of SpikingBrain to the MQL5 environment. It was important to demonstrate how theoretical principles can take shape in specific algorithms ready to work with real trading data. We began by implementing a basic component that converts a continuous signal into a discrete spike pulse. This paved the way from an abstract idea to a tool capable of operating within a popular trading platform. This work marked the first step toward building a fully fledged system in which the framework authors’ innovations are put into practice in a live market environment.
We have set the direction and laid the groundwork; now we'll move on to a more in-depth analysis of the model's architecture and operating principles.
Discussion of Implementation
Before moving on to the actual implementation of the proposed approaches, it is worth outlining the key principles of our work. The framework's authors use pre-trained models based on the standard Transformer, and this is an important detail. This approach allows us to avoid spending resources on training from scratch and instead build directly on the accumulated experience of architectures that have proven their effectiveness. Consequently, in our implementation, we can also utilize ready-made algorithms and modules, which makes the process more reliable and flexible.
The concept of hybrid attention modules is particularly interesting. The authors propose combining different architectural approaches within a single system — for example, using linear, windowed, or full attention in parallel. In some ways, this resembles the principles implemented in the Extralonger framework, where global-local attention has become the basis for a deeper analysis of temporal and spatial relationships.
Nevertheless, financial applications have their own unique characteristics, which require careful consideration when choosing an architecture. Thus, windowed attention — which has proven effective in a number of applications — is by no means always the optimal tool for analyzing market time series. On the one hand, the latest quote values do reflect the emerging trend, but on the other hand, this is where the highest level of noise is concentrated. The most significant signals influencing long-term trends are by no means always found in the most recent data points of a time series. Often, they are spread out over a greater distance and manifest as sparse dependencies; simply expanding the analysis window does not always solve the problem. The wider the window, the more computational resources are spent on processing noise rather than identifying useful patterns. As a result, the computational load on the model increases, while the quality of the forecast does not necessarily improve.
Therefore, when developing models for financial markets, it is preferable to use more adaptive attention mechanisms. Their task is to filter out the unnecessary and identify rare but truly valuable signals. It is precisely these signals that often prove decisive in shifting market trends and enable models to generate forecasts of practical significance.
In this regard, the idea of graphon-based sparse attention looks especially interesting. This approach allows us not simply to mechanically expand the analysis window, but to structure the relationships between elements in the time series, thereby identifying key dependencies in a more efficient manner. Essentially, graphons define a framework over which attention extends only in those directions where there are actual correlations and meaningful signals. This reduces the load on the model and allows resources to be focused on the data areas that really matter.
Moreover, the use of graphons fits seamlessly with the framework authors' vision of implementing algorithms such as Mixture of Experts (MoE). The sparse attention structure and the modular MoE architecture reinforce each other. Graphons define interaction paths, while MoE allows for the selective activation of the most appropriate expert blocks. The result is a system that can potentially handle noisy financial time series more effectively and adapt flexibly to the nature of the current market.
In fact, it is precisely the idea of combining biologically plausible signal-processing mechanisms with the architectural flexibility of modern neural networks that lies at the heart of SpikingBrain. The framework's authors propose an original solution that goes beyond traditional recurrent or Transformer-based architectures. They take a spiking neuron model as the basis, in which information is transmitted via discrete pulses, and learning is linked to the dynamics of interactions over time. This approach is more in line with the way the human brain works and opens up possibilities for more nuanced and efficient data processing.
In addition, the authors' proposed interpretation of the linear attention module is of particular interest. In one of our previous projects, we had already implemented a linear attention object and are very familiar with its capabilities. However, the framework's authors proposed a more flexible option. Linear attention is implemented as a recurrent block with added gates, which allows the scope of attention to be significantly expanded without a substantial increase in computational cost.
In this version, the attention state is computed recurrently. Each step accumulates information from the previous state, weighting it using gates, and adds new interactions between keys and values.
![]()
![]()
where St is the accumulated state matrix, gt is the gate vector, kt and vt are the keys and values at the current step, and qt is the query. This structure makes it possible to retain information about long-range dependencies while simultaneously controlling its relevance through gates.
From a financial analysis perspective, this solution is particularly valuable. It makes it possible to account for events spread out over time without overloading the model with unnecessary signals and noise. When combined with graphon-based sparse attention and MoE architecture, recurrent linear attention becomes a powerful tool for generating adaptive forecasts capable of identifying key trends and responding to rapidly changing market conditions.
The authors of the SpikingBrain framework propose two main approaches to using hybrid attention: sequential and parallel. The sequential variant involves the use of a stack of attention layers with different architectures, with blocks of the MoE type inserted between them to serve as FeedForward modules. This approach allows the model to process signals of different types sequentially, while maintaining structural flexibility and the ability to selectively activate experts.
Parallel mixing, on the other hand, involves the simultaneous use of different attention blocks within a single layer. Their combined result is then sent to a single MoE block for processing, which ensures the integration of information from different perspectives. Moreover, even in a parallel architecture, it is possible to combine different attention modules between the model's layers. The model achieves maximum flexibility, allowing the attention structure to be adapted to the specific characteristics of the data.
From the perspective of financial markets, this approach appears particularly attractive. Different information sources, time horizons, and signal types can be processed simultaneously, while the system remains manageable and efficient. When combined with graphon-based sparse attention and recurrent linear attention blocks, hybrid modes make it possible to build complex models capable of capturing both local noisy fluctuations and global market trends.
Linear Attention Module with Gates
Today, we will begin our practical work by implementing a recurrent linear attention block with gates, which we will implement as a new object called CNeuronGateLineAttention. This object is created as a specialized module within the SpikingBrain architecture and inherits its base functionality from CNeuronSpikeConv. This inheritance makes it easy to integrate the object into the existing infrastructure of spiking neural objects, take advantage of spiking dynamics and recurrent information accumulation, and use parallel computation via OpenCL to accelerate data processing.
This solution combines the best of both worlds. Within the object, standard continuous-valued multi-head linear attention is used, ensuring the accurate extraction of key signals and long-range dependencies. At the same time, the parent-class algorithms reduce the dimensionality of the multi-head attention output to the original data dimensionality and convert the resulting values into discrete spike pulses. Thus, the model retains the accuracy of continuous analysis while also gaining the ability to operate within event-driven spiking dynamics, which is particularly important for processing noisy and sparse financial time series.
The structure of the new object is shown below.
class CNeuronGateLineAttention : public CNeuronSpikeConv { protected: uint iDimensionK; uint iVariables; uint iHeads; //--- CNeuronConvOCL cQKV; CNeuronBaseOCL cQ; CNeuronBaseOCL cK; CNeuronBaseOCL cV; CNeuronBaseOCL cKV; CNeuronConvOCL cGate; CNeuronBaseOCL cState; CNeuronBaseOCL cMHAttention; //--- virtual bool feedForward(CNeuronBaseOCL *NeuronOCL) override; virtual bool updateInputWeights(CNeuronBaseOCL *NeuronOCL) override; virtual bool calcInputGradients(CNeuronBaseOCL *NeuronOCL) override; public: CNeuronGateLineAttention(void) {}; ~CNeuronGateLineAttention(void) {}; //--- virtual bool Init(uint numOutputs, uint myIndex, COpenCLMy *open_cl, uint variables, uint dimension, uint dimension_k, uint heads, ENUM_OPTIMIZATION optimization_type, uint batch); //--- virtual int Type(void) override const { return defNeuroneGateLineAttention; } //--- methods for working with files virtual bool Save(int const file_handle) override; virtual bool Load(int const file_handle) override; virtual void SetOpenCL(COpenCLMy *obj) override; virtual bool Clear(void) override; virtual void SetActivationFunction(ENUM_ACTIVATION value) override { }; //--- virtual bool WeightsUpdate(CNeuronBaseOCL *source, float tau) override; virtual uint GetVariables(void) const { return iVariables; } };
The object contains variables for storing the key parameters of its architecture:
- iDimensionK specifies the dimensionality of the keys, which determines the amount of information stored in each attention head;
- iVariables stores the number of variables being analyzed for signal processing;
- iHeads specifies the number of parallel attention heads, which enables multi-channel analysis of a time series.
Queries, keys, and values are processed using a combination of CNeuronBaseOCL and CNeuronConvOCL objects. The cQKV object plays a special role here, as it simultaneously generates all the components of multi-head attention. The unified result is then split into three tensors (q, k, and v) to perform the mathematical operations involved in attention. This approach allows all necessary components to be generated simultaneously, speeds up GPU computations, and reduces the load on the central processor, while maintaining the accuracy of multi-head analysis.
All internal components of our CNeuronGateLineAttention object are declared statically, so the class's constructor and destructor remain empty. The block is fully initialized in the Init method, which sets the object's operating parameters, connects the OpenCL compute core, and configures the internal components.
bool CNeuronGateLineAttention::Init(uint numOutputs, uint myIndex, COpenCLMy *open_cl, uint variables, uint dimension, uint dimension_k, uint heads, ENUM_OPTIMIZATION optimization_type, uint batch) { if(!CNeuronSpikeConv::Init(numOutputs, myIndex, open_cl, heads * dimension_k, heads * dimension_k, dimension, 1, variables, optimization_type, batch)) return false;
The method's algorithm begins by passing control to the parent class, CNeuronSpikeConv. This defines the basic structure of the object. This approach ensures consistency in configuration and allows for the use of proven processing and parallel computing algorithms that have already been implemented in the parent class.
In addition, as mentioned earlier, the parent class provides mechanisms for combining the results of multi-head attention into a single final signal, which is reflected in the passed parameters and facilitates the integration of a new object into the model’s overall architecture.
After successfully completing the parent initialization, we save the object's specific parameters: key dimensionality iDimensionK, the number of attention heads iHeads, and the number of variables being analyzed iVariables.
iDimensionK = dimension_k; iHeads = heads; iVariables = variables;
These parameters determine the scale of multi-head attention and the structure of the tensors that will accumulate and transmit information within the block.
Next, we move on to initializing the internal components. The first object we initialize is cQKV. It simultaneously generates all three components of multi-head attention — queries q, keys k, and values v.
int index = 0; if(!cQKV.Init(0, index, OpenCL, dimension, dimension, 3*iHeads*iDimensionK, 1, iVariables, optimization, iBatch)) return false; cQKV.SetActivationFunction(SIGMOID);
This approach speeds up computations, efficiently allocates GPU resources, and ensures the accurate construction of all components of multi-head attention. For cQKV, a sigmoid activation function is specified, which smoothly modulates the input signal and prepares it for recurrent processing.
Next, the cQ, cK, and cV objects are initialized.
index++; if(!cQ.Init(0, index, OpenCL, iHeads * iDimensionK * iVariables, optimization, iBatch)) return false; cQ.SetActivationFunction(None); index++; if(!cK.Init(0, index, OpenCL, cQ.Neurons(), optimization, iBatch)) return false; cK.SetActivationFunction(None); index++; if(!cV.Init(0, index, OpenCL, cQ.Neurons(), optimization, iBatch)) return false; cV.SetActivationFunction(None);
They are designed to store the corresponding tensors and are used in matrix multiplication operations. Activation functions are not used here.
The next key component is cKV. It is designed to store the results of matrix multiplication of the key and value tensors.
index++; if(!cKV.Init(0, index, OpenCL, cQ.Neurons()*iDimensionK, optimization, iBatch)) return false; cKV.SetActivationFunction(None);
As a result, the object accumulates the information needed to compute attention and passes it on to subsequent processing stages.
The cGate object is responsible for implementing the gating component of recurrent attention. Using a sigmoid activation function, it dynamically controls which part of the previous state is retained and how it is combined with the current input data.
index++; if(!cGate.Init(0, index, OpenCL, dimension, dimension, iDimensionK * iHeads, 1, iVariables, optimization, iBatch)) return false; cGate.SetActivationFunction(SIGMOID);
This is critically important for financial time series, where signals can be sparse and noisy: the gating mechanism makes it possible to filter out unnecessary fluctuations while retaining meaningful information from previous steps.
The cState object accumulates the previous attention state, enabling recurrent processing. It stores accumulated data on key dependencies over time, which are then used to generate the block's final output.
index++; if(!cState.Init(0, index, OpenCL, cKV.Neurons(), optimization, iBatch)) return false; cState.SetActivationFunction(None);
In combination with cKV and cGate, this makes it possible to implement recurrent linear attention with gates, which can account for both local short-term market fluctuations and global trends, ensuring the accurate identification of meaningful signals.
cMHAttention completes the internal configuration. It combines various attention channels and generates the final signal for subsequent computations.
index++; if(!cMHAttention.Init(0, index, OpenCL, cQ.Neurons(), optimization, iBatch)) return false; cMHAttention.SetActivationFunction(None); //--- return true; }
The Init method thoroughly checks whether each component was initialized successfully and returns false if any step fails. Thus, a fully operational object is created that can efficiently process financial time series, combining the precision of continuous multi-head attention with the advantages of discrete spiking dynamics and recurrent information accumulation.
Once the object has been initialized, the next step is to organize the forward pass. This process is implemented in the feedForward method, which handles the sequential processing of input data, recurrent state accumulation, and the generation of multi-head attention results.
bool CNeuronGateLineAttention::feedForward(CNeuronBaseOCL *NeuronOCL) { if(!cQKV.FeedForward(NeuronOCL)) return false;
In the first step, the method of the same name of the cQKV object is called. Here, the combined query q, key k, and value v tensors are created for multi-head attention. This approach speeds up computations because all three entities are generated simultaneously, and the subsequent separation into individual tensors allows standard attention operations to be performed.
Next, the FeedForward method of the cGate object is called, generating the gate signals.
if(!cGate.FeedForward(NeuronOCL)) return false;
The gating mechanism determines which part of the previous state is retained and which part is updated with the current input data. This is especially important for financial time series, where signals are sparse and prone to noise: gates help filter out noisy fluctuations and preserve meaningful patterns.
The next step is to split the combined cQKV output into three separate tensors using the DeConcat function.
if(!DeConcat(cQ.getOutput(), cK.getOutput(), cV.getOutput(), cQKV.getOutput(), iDimensionK, iDimensionK, iDimensionK, iHeads * iVariables)) return false;
After that, the keys are matrix-multiplied by the values using MatMul. The result is stored in the cKV object, which accumulates information for recurrent state accumulation, ensuring efficient data storage for subsequent computations.
if(!MatMul(cK.getOutput(), cV.getOutput(), cKV.getOutput(), iDimensionK, 1, iDimensionK, iHeads * iVariables, true)) return false;
Recurrent state updates are performed in the cState object. First, SwapOutputs is called to save the values from the previous forward pass.
if(!cState.SwapOutputs()) return false; if(!DiagMatMul(cGate.getOutput(), cState.getPrevOutput(), cState.getOutput(), iDimensionK, iDimensionK, iHeads * iVariables, cState.Activation())) return false; if(!SumAndNormilize(cState.getOutput(), cKV.getOutput(), cState.getOutput(), iDimensionK, false, 0, 0, 0, 1)) return false;
Next, DiagMatMul is applied using the gate outputs and the previous state. In this step, the block's new current state is formed. After that, the result is summed with the results from cKV. These operations isolate meaningful signals and suppress noise, ensuring that the information is accurately represented for the next step.
The final step is to perform matrix multiplication of the query cQ by the updated state cState, which produces the final output of multi-head attention in the cMHAttention object.
if(!MatMul(cQ.getOutput(), cState.getOutput(), cMHAttention.getOutput(), 1, iDimensionK, iDimensionK, iVariables * iHeads, true)) return false;
This result is passed to the parent class CNeuronSpikeConv to generate discrete spike pulses.
if(!CNeuronSpikeConv::FeedForward(cMHAttention.AsObject())) return false; //--- return true; }
As a result, continuous attention calculations are transformed into events that can be used in spiking dynamics and subsequent forecasting of market trends.
Taken together, the feedForward method creates a coherent data flow: tensor generation, gating regulation, recurrent state accumulation, and multi-head attention aggregation. This approach enables the model to work effectively with financial time series, taking into account short-term fluctuations and long-range dependencies, and to generate accurate signals for forecasting market trends.
Here, we have deliberately chosen not to include the residual connections typical of the standard Transformer architecture. This is because we plan to use several attention modules in parallel. The presence of a separate residual connection in each of them could shift the emphasis toward the input data and weaken the influence of the extracted signals. To eliminate this effect, we provide a single shared path for residual connections among all attention modules within a single layer of the model, which helps maintain a balance between analyzing the raw data and learning from the identified patterns.
However, the forward pass is only half the battle. To train the model, it is necessary to correctly distribute the prediction error among all components involved in the process according to their influence on the final result. In the CNeuronGateLineAttention object, this mechanism is implemented in the calcInputGradients method, which performs backpropagation and computes gradients for all internal components.
bool CNeuronGateLineAttention::calcInputGradients(CNeuronBaseOCL *NeuronOCL) { if(!NeuronOCL) return false; //--- if(!CNeuronSpikeConv::calcInputGradients(cMHAttention.AsObject())) return false;
In the first step, the method verifies that the pointer to the NeuronOCL source data object is valid. Next, the parent class's method with the same name is called, passing a pointer to the cMHAttention object, in order to propagate the gradients received from subsequent model objects to the multi-head attention mechanism, distributing each head’s contribution to the overall result.
The next step is to compute the gradients for the query cQ and the accumulated state cState. The MatMulGrad method distributes the error proportionally to the influence of each input signal on the final result of multi-head attention.
if(!MatMulGrad(cQ.getOutput(), cQ.getGradient(), cState.getOutput(), cState.getGradient(), cMHAttention.getGradient(), 1, iDimensionK, iDimensionK, iVariables * iHeads, true)) return false;
Next, the gradients are passed to the cKV object, adjusting the values according to the activation function.
if(!DeActivation(cKV.getOutput(), cKV.getGradient(), cState.getGradient(), cKV.Activation())) return false;
Special attention is given to the gating component. The DiagMatMulGrad method distributes the error between the gate output cGate and the previous state cState.
if(!DiagMatMulGrad(cGate.getOutput(), cGate.getGradient(), cState.getPrevOutput(), cKV.getPrevOutput(), cState.getGradient(), iDimensionK, iDimensionK, iHeads * iVariables)) return false; Deactivation(cGate)
After that, deactivation of the gating block is invoked to account for the activation function and correctly adjust the gradients.
Gradients for the keys cK and values cV are computed using MatMulGrad.
if(!MatMulGrad(cK.getOutput(), cK.getGradient(), cV.getOutput(), cV.getGradient(), cKV.getGradient(), iDimensionK, 1, iDimensionK, iHeads * iVariables, true)) return false;
Then, all the gradients (cQ, cK, cV) are recombined into a single tensor, cQKV, using the Concat function.
if(!Concat(cQ.getGradient(), cK.getGradient(), cV.getGradient(), cQKV.getGradient(), iDimensionK, iDimensionK, iDimensionK, iHeads * iVariables)) return false; Deactivation(cQKV)
This allows us to maintain the consistency of the gradients and correctly distribute the error among the components of multi-head attention.
Next, we need to pass the error gradients back to the input data level. Here, we pass values along two data streams. First, we pass the values from the cQKV object.
if(!NeuronOCL.CalcHiddenGradients(cQKV.AsObject())) return false;
After that, the pointer to the gradient buffer is temporarily replaced to preserve the previously obtained values.
CBufferFloat* temp = NeuronOCL.getGradient(); if(!NeuronOCL.SetGradient(PrevOutput, false)) return false;
Data is then transferred from the gating object cGate, followed by summing the values from the two information streams.
if(!NeuronOCL.CalcHiddenGradients(cGate.AsObject())) return false; if(!SumAndNormilize(PrevOutput, temp, temp, GetFilters(), false, 0, 0, 0, 1)) return false; if(!NeuronOCL.SetGradient(temp, false)) return false; //--- return true; }
The final step is to reset the pointers to the data buffers to their initial state so that they are ready for the weights to be updated during the parameter optimization phase.
Thus, the calcInputGradients method forms a complete flow of error propagation from the final multi-head attention to each internal component. This ensures proper training of recurrent linear attention with gates and allows the model to adapt effectively to financial time series, taking into account both local fluctuations and global market trends.
The final step in working with the CNeuronGateLineAttention object is optimizing the model parameters, which is implemented in the updateInputWeights method. This is where the weights of all internal components containing trainable parameters are updated.
bool CNeuronGateLineAttention::updateInputWeights(CNeuronBaseOCL *NeuronOCL) { if(!cQKV.UpdateInputWeights(NeuronOCL)) return false; if(!cGate.UpdateInputWeights(NeuronOCL)) return false; if(!CNeuronSpikeConv::updateInputWeights(cMHAttention.AsObject())) return false; //--- return true; }
The method sequentially passes control to the cQKV and cGate objects, updating their weights based on the accumulated gradients. Control is then passed to the parent class to update the parameters of the multi-head attention result aggregation block.
It is worth noting that there are few trainable parameters within CNeuronGateLineAttention, so the method simply performs a careful, sequential transfer of control between components. Nevertheless, this step is critically important: it is here that the model adapts to financial time series, adjusting the weights to accurately forecast trends and effectively recognize significant signals.
This concludes our work with the CNeuronGateLineAttention class. The complete code, along with all the methods, is included in the attachment to this article.
Today we did a great deal of painstaking work, systematically breaking down and implementing the key components of recurrent linear attention with gates. It is time to take a short break and let our thoughts settle. In the next article, we will continue our work and bring the implementation to its logical conclusion. We will train the model and test its effectiveness using real historical financial market data.
Conclusion
In this article, we analyzed and implemented in detail a recurrent linear attention block with gates within the CNeuronGateLineAttention object. Step by step, we worked our way from initialization and tensor generation to the forward pass, error backpropagation, and weight optimization, demonstrating how each component affects the processing of financial time series.
Particular attention was paid to the architectural decisions: gating regulation, recurrent state accumulation, the combination of multi-head attention, and a unified residual-connection path for all modules within the layer. All of this creates a mechanism capable of identifying meaningful signals, filtering out noise, and taking into account both short-term market fluctuations and long-term trends.
Today, we laid a solid foundation for the model's practical operation. In the next article, we will continue along this path, taking the implementation through the training and testing phases using real historical data. This will make it possible to assess how effective the proposed approaches are under real-world market conditions and how they help generate accurate forecasts.
Links
Programs used in the article
| # | Name | Type | Description |
|---|---|---|---|
| 1 | Study.mq5 | Expert Advisor | Expert Advisor for offline model training |
| 2 | StudyOnline.mq5 | Expert Advisor | Expert Advisor for online model training |
| 3 | Test.mq5 | Expert Advisor | Expert Advisor for model testing |
| 4 | Trajectory.mqh | Class library | Structure for describing the system state and model architecture |
| 5 | NeuroNet.mqh | Class library | Class library for creating a neural network |
| 6 | NeuroNet.cl | Library | Code library for an OpenCL program |
Translated from Russian by MetaQuotes Ltd.
Original article: https://www.mql5.com/ru/articles/19762
Warning: All rights to these materials are reserved by MetaQuotes Ltd. Copying or reprinting of these materials in whole or in part is prohibited.
This article was written by a user of the site and reflects their personal views. MetaQuotes Ltd is not responsible for the accuracy of the information presented, nor for any consequences resulting from the use of the solutions, strategies or recommendations described.
Building AI-Powered Trading Systems in MQL5 (Part 12): Giving the Assistant Chart Vision and Tool Access
Encoding Candlestick Pattern (Part 6): Developing the Encoded Sequence Indicator
Trade Entry Timing Accuracy Analyzer in MQL5
Regime Discovery by Structure: Implementing Toeplitz Inverse Covariance Clustering (TICC)
- Free trading apps
- Over 8,000 signals for copying
- Economic news for exploring financial markets
You agree to website policy and terms of use
The article "Neural Networks in Trading: From Transformers to Spiking Neurons (Basic Components)" has been published:
Author: Dmitry Gizlyk