Memory alignment issue with C++20 DLL bridge into MT5 MQL5 EA (Lock-free SPSC context)

 

Hi everyone,

I am currently optimizing a ultra-low latency ingestion layout between an external bare-metal C++20 HFT execution core and an MT5 EA via a custom native DLL bridge using lock-free Single Producer Single Consumer (SPSC) ring buffers.

During volatile gold tick regimes (XAUUSD ), the core engine handles the internal wire baseline comfortably at 150ns - 247ns using flat contiguous cache arrays. It enforces an in-memory Anti-Trap Interceptor to process Level 3 data structures within absolute O(1) time complexity.

However, when scaling the packet synchronization loops within a 1ms timer gate down to an exact 20-pips execution boundary , I am noticing occasional serialization drag at the terminal interface. The target metrics are structured as follows:
  • Memory structures: Aligned to 64-byte L1 CPU cache boundaries via alignas(64) .
  • Data ingestion: Zero-allocation arrays parsing Order Flow Imbalance (OFI) metrics.

Has anyone successfully eliminated standard Windows OS thread-switching latency inside an MQL5 DLL import layer when passing active pointers from contiguous C++ memory slots?

(Note: If any senior algorithmic engineer or institutional developer wants to cross-verify the raw .html interface simulator or review the complete cache block architecture to locate the pipeline drag, let me know. I can share the repository manual setup via private message).
 

Sorry but.. Question is not making sense ..

what is internal wire baseline comfortably at 150ns - 247ns?

what is Anti-Trap Interceptor / absolute O(1) Level 3 structures?


Thread switching can usually take 1 to 3 or 5 ms in micro processors, NS no way..

 
Selim Hossain:

Hi everyone,

I am currently optimizing a ultra-low latency ingestion layout between an external bare-metal C++20 HFT execution core and an MT5 EA via a custom native DLL bridge using lock-free Single Producer Single Consumer (SPSC) ring buffers.

During volatile gold tick regimes (XAUUSD ), the core engine handles the internal wire baseline comfortably at 150ns - 247ns using flat contiguous cache arrays. It enforces an in-memory Anti-Trap Interceptor to process Level 3 data structures within absolute O(1) time complexity.

However, when scaling the packet synchronization loops within a 1ms timer gate down to an exact 20-pips execution boundary , I am noticing occasional serialization drag at the terminal interface. The target metrics are structured as follows:
  • Memory structures: Aligned to 64-byte L1 CPU cache boundaries via alignas(64) .
  • Data ingestion: Zero-allocation arrays parsing Order Flow Imbalance (OFI) metrics.

Has anyone successfully eliminated standard Windows OS thread-switching latency inside an MQL5 DLL import layer when passing active pointers from contiguous C++ memory slots?

(Note: If any senior algorithmic engineer or institutional developer wants to cross-verify the raw .html interface simulator or review the complete cache block architecture to locate the pipeline drag, let me know. I can share the repository manual setup via private message).

The only way to reliably remove OS interference is by manual core assignment. Meaning, you need d to evacuate the core (both, hyperthreaded and base core) from all threads in the system via affinity mask, then, put your thread onto that core.

But in reality I would ask, why are you trying to do that with MQ as your data source in the first place? - Your approach will not grant you any edge over others. Its nice to see you did it, but for what?

If you are working on a direct connect to an exchange match engine, you would not need/use MQ software in the first place.

It seems a bit like mounting an F1 tire to a bike.
 
Farrukh Aleem #:

Sorry but.. Question is not making sense ..

what is internal wire baseline comfortably at 150ns - 247ns?

what is Anti-Trap Interceptor / absolute O(1) Level 3 structures?


Thread switching can usually take 1 to 3 or 5 ms in micro processors, NS no way..


Thread switching takes about 30 to 50 CPU cycles, that's in the range of nanoseconds. If thread switching would take milliseconds, no game would ever run smooth.
 
Dominik Egert #:

Thread switching takes about 30 to 50 CPU cycles, that's in the range of nanoseconds. If thread switching would take milliseconds, no game would ever run smooth.

He meant microseconds I guess, not milliseconds. 

30 to 50 cycles is for a user-level thread switching, I am not sure that can apply here in a setup implying MT5, a DLL and C++  ("standard Windows OS thread-switching latency"), how do you come to that conclusion ?

 
Selim Hossain:
I am noticing occasional serialization drag at the terminal interface.

What does that mean exactly ? We can only guess you have a bottleneck...somewhere. It's rather vague indication of your actual problem.

Selim Hossain:
Has anyone successfully eliminated standard Windows OS thread-switching latency inside an MQL5 DLL import layer when passing active pointers from contiguous C++ memory slots?
How do you come to the conclusion OS thread-switching is the (only) problem ?
 
Dominik Egert #:

Thread switching takes about 30 to 50 CPU cycles, that's in the range of nanoseconds. If thread switching would take milliseconds, no game would ever run smooth.

Windows thread switching (context switching) latency typically takes between 1 to 5 microseconds on modern consumer CPUs, though the overall scheduler interval (quantum) dictates when switches happen globally every 10 to 15 milliseconds

  • Context Switch Time: The actual CPU cost to save one thread's state and load another takes roughly 1,000 to 5,000 nanoseconds under normal conditions.
  • The Quantum: Windows allocates execution slices (quanta) to threads—typically 2 clock intervals (10–15 ms) on client operating systems like Windows 10 and 11.
  • Forced Switching: If a higher-priority thread becomes ready, or a thread exhausts its quantum/blocks on I/O, the Windows scheduler triggers a context switch immediately.
 
Alain Verleyen #:

He meant microseconds I guess, not milliseconds. 

30 to 50 cycles is for a user-level thread switching, I am not sure that can apply here in a setup implying MT5, a DLL and C++  ("standard Windows OS thread-switching latency"), how do you come to that conclusion ?


I am referring to thread switching within the same process frame/descriptor. A DLL is loaded into the process of MT5, an expert is a thread within this process, a thread created within the DLL will be a client of MT5 process descriptor.

Therefore I concluded the switching of that thread would require the saving and restoration of the main registers and the program counter.

Therefore, roughly estimated around 30 to 50 cycles for copying these data points to memory, and restoring the others threads state.

I took into account the 16 main x86 registers, and the copying of those. Which brought me to an estimate of 30 to 50 cycles.

There are a lot of other factors, which also would need consideration.
 
Dominik Egert #:

I am referring to thread switching within the same process frame/descriptor. A DLL is loaded into the process of MT5, an expert is a thread within this process, a thread created within the DLL will be a client of MT5 process descriptor.

Therefore I concluded the switching of that thread would require the saving and restoration of the main registers and the program counter.

Therefore, roughly estimated around 30 to 50 cycles for copying these data points to memory, and restoring the others threads state.

I took into account the 16 main x86 registers, and the copying of those. Which brought me to an estimate of 30 to 50 cycles.

There are a lot of other factors, which also would need consideration.

I got you but from my understanding, with the information provided there is no way to know what thread mechanism is used. 

So it could be standard OS thread which require 1 to 5 µs or some user thread requiring 30 to 50 cycles. 

Additionally we don't even now what the actual problem is, are we sure the thread switching is the problem ? We can't state it from the provided information. Unless I misunderstood something (?).

Of course, I do agree on your general points about core assignment, and that it's strange to use such system with an MT5 backend. 

 
Alain Verleyen #:

I got you but from my understanding, with the information provided there is no way to know what thread mechanism is used. 

So it could be standard OS thread which require 1 to 5 µs or some user thread requiring 30 to 50 cycles. 

Additionally we don't even now what the actual problem is, are we sure the thread switching is the problem ? We can't state it from the provided information. Unless I misunderstood something (?).

Of course, I do agree on your general points about core assignment, and that it's strange to use such system with an MT5 backend. 


Yeah, missing details. We cannot conclude the issues source from what's given. Could be an interlocking problem on a cache line, a MMU page invalidation, or something as simple as an interrupt firing.

To many variables outside the scope of the actual program.