Architecture for Collective Trading Decisions by AI Agents
How to Make a Language Model Doubt Itself
When a trader looks at a chart and makes a decision, they never do so alone, even if there is no one else in the room. Several voices speak at once in their head. One notes that the price has broken up through the moving average from below and that momentum is positive. Another objects: the RSI is already at 68, the stochastic is in overbought territory, and the latest candle has a long upper shadow — someone is actively selling at these levels. A third says nothing at all about direction and says only this: the ATR is three times higher than normal today; it is a news-driven day, so any position right now is a gamble.
A professional trader knows how to listen to all three voices at the same time and weigh them against one another. A beginner hears only the first one — and loses money on what seemed obvious.
When we connect a large language model to MetaTrader 5 — which was covered in the previous article in this series, describing the Shtenco AI V17 architecture with a WebSocket server and the PRICES command — we are essentially replacing that entire internal dialogue with a single voice. The model receives data, processes it using its single system prompt, and returns a response. The answer may be right, or it may be wrong, but there is always only one — without doubt, without contradiction, and without weighing the alternatives. That is the problem — not a technical one, but an architectural one.
A language model that operates with a single system prompt such as “You are a professional trader; give clear signals” will inevitably tend to generate signals. It is optimized for the task it has been assigned. If it is told, “give buy or sell”, it will give buy or sell, even when the market is screaming, “Stop, this is not your moment.” A neutral hold in this configuration is effectively a losing outcome for a model trying to appear useful.
Why One Prompt Is Not Enough
If you look at how the best trading-desk teams at banks and hedge funds operate, they have long had a separation of roles that, at first glance, seems excessive. There is an analyst who builds the bullish case. There is another analyst whose job is to identify its weaknesses and formulate a bearish scenario. There is a risk manager who doesn’t care at all where the price is headed; their job is to answer the question, “Should we even be in the market right now?” And there is a senior portfolio manager who listens to everyone and makes the final decision, taking positions across the entire portfolio into account.
This is not bureaucracy or excessive caution. This is the only known way to combat what behavioral economists call “confirmation bias” — the tendency to seek confirmation of a decision that has already been made and to ignore conflicting signals. A person who comes up with a trading idea is psychologically incapable of critiquing it objectively. That is why the critic is a distinct role.
A language model suffers from the same problem, just in a different form. If its system prompt says “look for bullish signals,” it will find bullish signals even where there are not any. This is not a hallucination in the technical sense of the word; it is the normal operation of a transformer performing its assigned task. The solution is obvious: we need several instances of the model, each assigned different tasks and operating on the same data, along with a separate arbiter that delivers a verdict by weighing their disagreements rather than merely following consensus.
Debate Architecture
Four analysts, the same data
Four analysts. The same market data. Four fundamentally different system prompts — four different “lenses” through which the same grok-4-fast model looks at the same EURUSD candle.
Victor — the bull — is instructed to look for every possible bullish argument: moving-average alignment, the RSI exiting the oversold zone, positive momentum across multiple time horizons, a bounce off the lower Bollinger Band, and a volume spike on a bullish candle. If any of these are present, Victor is required to build an argument. He is an optimist by role, and this is not a weakness of the system but its intentional design.
Maria is the bear - the counterpart to Victor: overbought RSI, rejection from the upper Bollinger Band, a candlestick with a long upper shadow, price below the EMA55, negative momentum on higher timeframes while positive on shorter ones — this is always a sign of a weakening trend. Maria sees danger where Victor sees opportunity.
Alexey is a risk manager; he is not interested in the direction at all. He's focused on just one thing: whether it's safe to trade right now. Is the ATR above its historical average today? Are the Bollinger Bands unusually wide? If the candle body is three times the ATR, it means we're already too late; entering at this level means chasing the price. Alexey gives one of three verdicts: LOW RISK, MEDIUM RISK, or HIGH RISK — and never specifies the direction.
The arbiter — a judge — receives all three opinions and issues a final verdict based on strict rules hard-coded directly into its system prompt. Rule one: if Alexey says HIGH RISK, the final signal is always hold, regardless of what Victor and Maria say. No exceptions. Rule two: buy or sell is issued only if at least two out of the three support the same directional outcome. If Victor shouts buy, Maria shouts sell, and Alexey says MEDIUM RISK, this is uncertainty, and the honest response to uncertainty is hold.
ANALYST_PROMPTS = {
"bull": (
"You are VICTOR — an aggressive trend-following trader. "
"Your job is to find every possible BULLISH argument in the data. "
"Focus on: upward momentum, MA alignment (price above MA), "
"RSI coming out of oversold, bullish candle patterns, "
"Bollinger lower band bounces, volume surges on up-moves. "
"If even one bullish signal exists — argue for BUY. "
"Reply in 2-4 concise sentences in English. No JSON needed."
),
"bear": (
"You are MARIA — a skeptical contrarian analyst. "
"Your job is to find every possible BEARISH argument in the data. "
"Focus on: downward momentum, price below MA, "
"RSI overbought or declining, bearish candle patterns, "
"Bollinger upper band rejection, negative momentum divergence. "
"If even one bearish signal exists — argue for SELL. "
"Reply in 2-4 concise sentences in English. No JSON needed."
),
"risk": (
"You are ALEXEI — a risk manager and volatility specialist. "
"Your job is NOT to predict direction, but to evaluate trade SAFETY. "
"Analyze: ATR level (high ATR = dangerous), "
"Bollinger Band width, StdDev levels, "
"candle body size vs ATR (large body = chasing, small = good entry). "
"Give a risk verdict: LOW RISK / MEDIUM RISK / HIGH RISK. "
"Explain why in 2-4 sentences. No JSON needed."
),
"judge": (
"You are the ARBITER — an impartial senior analyst. "
"You will receive three expert opinions: BULL, BEAR, RISK MANAGER. "
"Rules: "
"1) If RISK says HIGH RISK → default to hold regardless of direction. "
"2) Only give buy/sell if at least 2 of 3 analysts agree on direction. "
"3) If BULL and BEAR disagree and RISK is MEDIUM → pick stronger argument. "
"Reply ONLY with JSON, no markdown:\n"
'{"signal":"buy"|"sell"|"hold","comment":"reasoning up to 150 chars"}'
),
}A crucial detail: analysts respond in free-form text, without JSON. They are expected to provide reasoning, not a verdict. The verdict is the arbiter's responsibility, and only the arbiter's response is parsed as JSON. The judge's temperature is set to 0.2 — significantly lower than that of the three analysts, who operate at 0.7. Analysts need to be a little “creative”: they need to find arguments even where someone else would not have found any. The judge must be as deterministic as possible: given the same input, it must produce the same output. That's the difference between "thinking" and "deciding."
Fifteen indicators in pure NumPy
All four analysts receive the same market briefing — a structured text block assembled by the _build_market_brief() function. Fifteen technical indicators calculated locally, without TA-Lib or any other third-party technical analysis libraries.
def _build_market_brief(symbol: str, ind: dict) -> str: """Unified briefing — same for all analysts.""" return ( f"Symbol: {symbol} | Bars: {ind['n']}\n" f"Last candle: {ind['candle']}\n" f"Volume: {ind['volume']}\n\n" f"Moving averages:\n {ind['ma']}\n {ind['ema']}\n\n" f"Oscillators:\n {ind['rsi']}\n Stoch {ind['stoch']}\n\n" f"Volatility:\n {ind['atr']}\n {ind['stddev']}\n\n" f"Bollinger Bands: {ind['bb']}\n" f"Momentum: {ind['momentum']}" )
Moving averages — six values: MA5, MA10, MA20, MA50, MA100, and MA200, plus three exponential moving averages — EMA9, EMA21, and EMA55. For each of the three key MAs, the briefing explicitly indicates the price’s position relative to it — UP or DOWN — which allows the model to instantly assess the trend structure. RSI is provided for three periods: 7, 14, and 21 — this gives a picture of momentum across short, medium, and long horizons. The stochastic oscillator is provided in the classic K/D format with a period of 14. Volatility: two ATR values (14- and 21-period) and two StdDev values (10- and 20-period). Bollinger Bands with an explicit price position: UPPER, MID, or LOWER. Momentum across three horizons: 5, 10, and 20 candles. A description of the latest candlestick, including its body, wicks, and direction.
The `build_indicators()` function accepts optional `high`, `low`, and `vol` arrays. If the Expert Advisor passes only comma-separated closing prices, the `high` and `low` values are automatically copied from `close`. This ensures full backward compatibility with Expert Advisors written for V17.
Parallelism in the first phase
The main technical problem with a system that uses multiple analysts is obvious: if requests are sent to the model sequentially, the response time increases linearly. Each request takes 3–5 seconds. Three consecutive requests already take 9–15 seconds, plus one more for the judge. This is unacceptable for live trading.
The solution is to parallelize the first phase. Three analysts start at the same time via a ThreadPoolExecutor with three worker threads. Their requests are sent to the xAI API all at once, processed independently, and returned in the order in which they are completed. In practice, three analysts complete in the same amount of time that one would take — 3–5 seconds. The judge adds another 2–3 seconds. Total: 5–8 seconds for a full debate run on a single symbol.
def run_debate(prices: list[float], symbol: str) -> dict: close = np.array(prices, dtype=float) ind = build_indicators(close) brief = _build_market_brief(symbol, ind) # Phase 1: Three analysts working in parallel ───────────────── opinions: dict[str, str] = {} with ThreadPoolExecutor(max_workers=3) as pool: futures = { pool.submit(_analyst_call, role, brief): role for role in ["bull", "bear", "risk"] } for fut in as_completed(futures): role, opinion = fut.result() opinions[role] = opinion # Phase 2: Judge after all opinions have been received ────────── judge_raw = _judge_call(brief, opinions) # ... parsing JSON and generating a response
The judge is launched only after all three opinions have been collected — `as_completed()` blocks the loop from exiting until the last future has completed. This is not a parallel task by its very nature: the arbiter must review all three before delivering a verdict.
Debate runs for each symbol are launched in a separate thread within `handle_client()`. If the Expert Advisor sends eight DEBATE commands nearly simultaneously, the server will launch eight independent debate runs — for a total of 24 parallel API calls in the first phase. This is fine for a paid API without strict rate limits, but it is important to keep this in mind when choosing a pricing plan.
The judge receives the input data, not just opinions
The judge receives not only three opinions but also the original market briefing — it sees exactly what the analysts saw, plus their arguments on top of that data.
def _judge_call(brief: str, opinions: dict) -> str: judge_input = ( f"MARKET DATA:\n{brief}\n\n" f"━━━ ANALYST OPINIONS ━━━\n\n" f"BULL (Victor):\n{opinions.get('bull','N/A')}\n\n" f"BEAR (Maria):\n{opinions.get('bear','N/A')}\n\n" f"RISK MANAGER (Alexei):\n{opinions.get('risk','N/A')}\n\n" f"━━━ YOUR TASK ━━━\n" f"Based on all three opinions above, deliver your FINAL verdict." ) messages = [ {"role": "system", "content": ANALYST_PROMPTS["judge"]}, {"role": "user", "content": judge_input}, ] # The judge is deterministic, temperature=0.2 return _call_api(messages, temperature=0.2, max_tokens=256, label="JUDGE")
A Hold signal issued by the arbiter when opinions conflict is not a weakness of the system. This is its most valuable signal. The ability to say "I don't know" is rare in any field. A trading system that knows when to remain silent in the face of uncertainty is worth more than one that always has something to say. That is precisely why the rule “HIGH RISK → hold without discussion” is embedded in the judge’s prompt as an absolute priority, rather than as a recommendation.
New Protocol: the DEBATE Command
The system is fully backward compatible with V17. All the old commands — PRICES, CHAT, CLEAR, STOP — work unchanged. The new DEBATE command follows the same logic as PRICES, but returns an extended response. Adding debate support to the MQL5 Expert Advisor is literally just one substitution in the command-construction line:
// Standard mode (V17, one analyst): string cmd = "PRICES:EURUSD:" + csv; // Debate mode (four analysts + judge): string cmd = "DEBATE:EURUSD:" + csv;
The response to the DEBATE command contains the same required minimum — the `signal` and `comment` fields — plus an extended `debate` block with the opinions of all four participants. An Expert Advisor that cannot parse the extended block simply ignores it and works with `signal` and `comment` as before.
{
"signal": "buy",
"comment": "2/3 bullish: MA alignment confirmed, MEDIUM risk",
"debate": {
"bull": "Price above all MAs, EMA9 crossing EMA21 upward...",
"bear": "RSI7 at 67 approaching overbought, upper shadow...",
"risk": "MEDIUM RISK: ATR14 normal, BB width stable...",
"judge": "{\"signal\":\"buy\",\"comment\":\"Bull/Risk consensus\"}"
}
}
The `debate.judge` field contains the arbiter’s raw JSON response — this is useful for debugging and logging. If you're building a system that uses SQLite for storage, it makes sense to store not only the final `signal` but also the entire `debate` field. After a week of live trading, you will have a dataset showing which patterns of analyst disagreement made the final signal profitable and which did not.
What the Expert Advisor Sees in the Log
One of the main practical advantages of the debate system is that the transparency of each decision becomes multidimensional. Previously, the Expert Advisor wrote something like this to the log: Signal [EURUSD]: buy | MA20 broken to the upside, RSI=58, positive momentum. One line, one opinion.
Now the MetaTrader log shows the full picture of the discussion — and this is not just a longer log; it is a fundamentally different level of understanding of why the system made a particular decision.
[12:34:11] DEBATE [EURUSD] started --- 60 bars [12:34:11] ↳ Analyst [BULL] thinking... [12:34:11] ↳ Analyst [BEAR] thinking... [12:34:11] ↳ Analyst [RISK] thinking... [12:34:14] ✓ [BULL]: Price above MA20/50/200, EMA9 crossing EMA21... [12:34:15] ✓ [RISK]: MEDIUM RISK: ATR14 within normal range... [12:34:16] ✓ [BEAR]: RSI7 at 67, upper shadow 40% of body... [12:34:16] ↳ JUDGE deliberating... [12:34:18] ✓ [JUDGE]: {"signal":"buy","comment":"Bull/Risk consensus"} [12:34:18] ══ FINAL: BUY | Bull/Risk consensus, bear shadow noted
Looking at these lines, a trader sees not just a signal — they see the context behind the signal. They see that the bear has spotted an alarming shadow. They see that the risk manager deemed the situation manageable. They see that the arbiter weighed everything, taking all three positions into account. This represents a fundamentally different level of confidence in the signal than "the model said 'buy.'"
Particularly valuable are instances where the system issues a "hold" not because of neutral indicators, but because Victor and Maria presented directly opposing arguments of equal weight. It is not that "nothing is happening" — it is that "we honestly do not know, and the right answer is not to trade." The ability to recognize such moments and refrain from making a trade is, in and of itself, more valuable than any signal.
Deployment and Compatibility
If you already have a server up and running from a previous article in this series, the transition takes just a few minutes. The llm_server_grok_debate.py file is a direct replacement for llm_server_grok.py: same port (8971), same WebSocket protocol, same commands — plus a new command, DEBATE. The dependencies remain the same: Python 3.8 or newer, and the requests and numpy packages. No additional libraries.
A sensible approach to deployment: use DEBATE to make final trading decisions — opening and closing positions — and reserve PRICES for quickly monitoring market conditions. Debate runs take 5–8 seconds; for an "open position" event, this is acceptable. For a quick overview of the eight currency pairs at the start of the trading session, the standard mode is sufficient.
One thing to keep in mind when working with multiple symbols at the same time: each call to DEBATE runs in a separate thread. If the Expert Advisor sends eight DEBATE commands almost simultaneously, the server will launch eight independent debate runs— for a total of 24 parallel API requests in the first phase. This is normal for a paid API without strict rate limits, but it is important to keep this in mind when choosing a pricing plan.
To manage expenses, use the same control as in V17: the InpAnalysisBars parameter in the Expert Advisor. When set to 3 on the M15 timeframe, there is one full set of debate runs every 45 minutes per symbol. Across eight currency pairs — approximately 256 debate runs per trading day. At grok-4-fast pricing, that comes out to a few cents a day — the cost of a cup of coffee per month.
On the Nature of Consensus and Its Limits
It is worth being honest about what the debate system is not. It is not an ensemble in the strict machine-learning sense — we are not averaging the probabilities of several independently trained models. All four analysts are the same grok-4-fast model with different system prompts. Their "independence" is the independence of perspectives, not the independence of the neural network weights.
This is both a limitation and an advantage. Limitation: all four "voices" share the same systematic errors inherent in the base model. If grok-4-fast performs poorly on a specific pattern, all four analysts will perform poorly on it. Advantage: different system prompts create genuine variability in how the same data is interpreted. A bullish prompt activates certain attention patterns in the transformer, while a bearish prompt activates others.
The next natural step, which suggests itself, is to give the system memory. Record each decision in SQLite along with what happened to the price afterward. In a month, you will have a dataset showing which configurations made the analysts' consensus correct, and which ones made it systematically wrong. This is the foundation for iteratively improving prompts — not based on intuition, but on data. That is exactly what the next article in this series will focus on.
On the Results of the “Right out of the Gate” Backtest
Before we break down the numbers, the main point must be stated: a backtest on historical data is not proof that the system works; it is its initial medical checkup. A patient may pass it with flying colors and still get sick. But if the patient has not even passed the medical exam, there is no point in discussing it further.
The system passed.
The MetaTrader 5 tester ran the Expert Advisor on EURUSD on a 30-minute timeframe in "All Ticks" simulation mode from February 2, 2026, to March 20, 2026. I have included the settings in the .set file attached to this article:

Unfortunately, because of the very long debate runs (which you can configure by adjusting the models or changing the length of the context window, the message history, and the temperature), testing the model takes a very long time.

The Sharpe ratio almost reached the acceptable threshold of 1.0; the percentage of profitable trades (66%) is good, but the average profit per winning trade is lower than the average loss per losing trade.
Next, it will be interesting to explore a model of collective interaction, competition, rivalry, and debate among various agents combined into clockwork-like systems that emulate the operations of a real hedge fund (the work of analysts, managers, risk managers, macro analysts, etc.) — we will cover all of this in future posts.
The next natural step, which suggests itself, is to give the system memory. Record each decision in SQLite along with what happened to the price afterward. In a month, you will have a dataset showing which configurations made the analysts' consensus correct, and which ones made it systematically wrong. This provides a foundation for iteratively improving prompts based on data rather than intuition. That is exactly what the next article in this series will focus on.
Conclusion
We started with a simple observation: one system prompt means one voice, and one voice in trading always means risk. Professional trading teams have long known this and structure their processes with a separation of roles — not because it is efficient in terms of headcount, but because it is the only known way to combat confirmation bias in financial decision-making.
The architecture described in this article applies this principle to the world of LLM trading. Three parallel analysts with deliberately different perspectives, plus a deterministic arbiter, are not complexity for complexity’s sake. This is an attempt to build into the system something that a single prompt, by definition, lacks: institutional skepticism.
Technically, the system is a direct extension of V17. The same Python server architecture serves as a bridge between MetaTrader 5 and the language model. The same WebSocket protocol, without any third-party dependencies. The same 15 technical indicators written entirely in NumPy. The only new additions are the run_debate() function and four system prompts. The rest of the code has not changed by a single line.
The result you get as output is not just "buy," "sell," or "hold." This is a well-reasoned position put forward by three experts with different professional backgrounds and a balanced verdict from an arbiter who is familiar with all the contradictions. Each decision reads like a professional analytical briefing, not a “black box” prediction.
All of this runs on a single machine, costs just a few cents per trading day, and is launched with a single command: python llm_server_grok_debate.py. The next step is up to you.
Translated from Russian by MetaQuotes Ltd.
Original article: https://www.mql5.com/ru/articles/21690
Warning: All rights to these materials are reserved by MetaQuotes Ltd. Copying or reprinting of these materials in whole or in part is prohibited.
This article was written by a user of the site and reflects their personal views. MetaQuotes Ltd is not responsible for the accuracy of the information presented, nor for any consequences resulting from the use of the solutions, strategies or recommendations described.
Self-Exciting Markets: Building a Hawkes Process from Scratch
From Deal History to Hazard Curves: Survival Analysis Applied To Strategies
Beyond REST and ZeroMQ: Building a gRPC/Protocol Buffers Bridge for Real-Time MetaTrader 5–Python Inference
Encoding Candlestick Patterns (Part 5): Expanding Taxonomy of Candlestick for General Pattern Frequency Analysis
- Free trading apps
- Over 8,000 signals for copying
- Economic news for exploring financial markets
You agree to website policy and terms of use