← All posts

FIX Lab and the engineering of FIX session recovery

FIX Lab shows the part of FIX most teams fail to test well, session recovery after message loss. The value is in the transcript: sequence numbers, resend handling, replay behavior, and the final state you can verify.

Market connectivity fails in small, repeatable ways long before it fails in dramatic ones. A sequence number drifts. A resend range arrives late. One side logs a gap while the other side thinks the session is healthy. If you work with FIX, you know the hard part is rarely sending a NewOrderSingle. The hard part is seeing session state clearly when messages go missing.

FIX Lab matters because it turns a usually hidden process into something you can inspect. You get your own FIX session. You lose a message on purpose. Then you watch the engines recover it. For a practitioner, this is the core lesson. Session reliability is a protocol behavior, not a promise. You verify it by observing sequence handling, recovery messages, and final state.

This kind of tool is useful beyond training. It exposes the narrow place where many trading systems break. Not at business logic. At session control. If your engine mishandles a gap, duplicates application messages after a resend, or advances sequence numbers in the wrong order, the bug often stays latent until a counterparty or venue behaves differently from your test harness. A controlled lab helps you inspect those edges without guessing.

What problem it addresses

FIX sessions look simple on a diagram. Logon. Heartbeats. Application flow. Logout. Real sessions are stricter. Each side tracks inbound and outbound sequence numbers. Each side has to detect missing messages and request retransmission. If one message is lost, the session needs to converge on a shared state again.

That is where implementation details matter.

A usable session lab addresses three practical problems:

  • It makes message loss observable.
  • It shows the recovery path, not only the failure.
  • It lets you compare what the protocol expects with what your engine does.

For many teams, the largest risk is false confidence from happy-path testing. If your test only proves message exchange under no loss, you have not learned much about session correctness. FIX recovery is where state machines, timers, and storage policy interact. A lab with intentional loss forces those paths to execute.

As you inspect the behavior, keep your focus narrow. You are not looking for a verdict on an engine or a venue. You are looking for signals. Did the peer detect the gap at the expected point. Did it issue the right recovery request. Did the resend fill the hole without corrupting later flow. Those are engineering questions with visible traces.

What to inspect in a lost-message recovery

When a message disappears, sequence discipline takes over. This is the first place to look.

Start with the sequence numbers around the missing message. You want to see the message before the gap, the first message after the gap, and the point where the receiver realizes it skipped a number. In a healthy recovery flow, the receiver should react to the out-of-sequence arrival by asking for the missing range.

Inspect these signals:

  • The inbound and outbound sequence numbers on both sides.
  • The exact point where the gap is detected.
  • The recovery request range.
  • The replayed messages and their sequence numbers.
  • The state after replay completes.

In practice, a gap recovery often includes a Resend Request and one of two replay styles. The sender either retransmits stored messages or uses gap fills where session rules permit it. The important part for verification is consistency. The receiver should end up with a coherent sequence stream. The sender should not improvise numbering or skip state transitions.

You should also watch for duplicate handling. A replayed application message is a normal event in a resend flow. Your system needs a policy for idempotency and for distinguishing replay from first delivery. If your downstream logic treats all replays as fresh business events, the session may recover while the order state corrupts.

This is where labs are useful. They isolate the session event from the rest of your stack. You can study the protocol before your own application code muddies the picture.

How to verify what the lab claims

A good demo is easy to admire and hard to trust unless you know how to test it. The claim here is specific. You lose a message, then the engines recover it. You should verify that claim from the transcript of session behavior.

Use a simple checklist:

  1. Establish the initial session state. Confirm sequence numbers begin where you expect.
  2. Trigger or observe the lost message event.
  3. Confirm the next higher sequence number arrives after the missing one.
  4. Check whether the receiver detects the gap, rather than silently accepting the stream.
  5. Inspect the recovery request. The requested range should match the hole.
  6. Inspect the sender response. It should replay or gap-fill according to session rules.
  7. Confirm the session returns to a stable sequence state.

The key here is reproducibility. If you repeat the exercise, the same class of event should produce the same class of recovery trace. If the behavior changes from run to run without a configuration change, look for hidden state, weak persistence, race conditions, or timer sensitivity.

Verification also means checking the boundaries. Lose an early message. Lose a later one. Observe whether the session reacts consistently. A fragile implementation sometimes works for a single-message gap in the middle of a stream but fails near logon, after a sequence reset, or across a reconnect.

If logs are visible, compare both sides. FIX bugs often hide in asymmetric interpretation. One side thinks it asked for 5 through 7. The other thinks it satisfied the request with a broader replay. Both sides continue, but your audit trail now disagrees. A reliable lab makes those mismatches obvious.

Where FIX session recovery commonly goes wrong

Most session bugs come from state, persistence, or assumptions.

One common failure is weak sequence persistence. If an engine stores sequence state late, or only in memory, a process restart changes behavior. After reconnect, the peer asks for a range your side no longer believes exists. That leads to confusing resend behavior and hard-to-reproduce support cases.

Another failure is incorrect resend logic. Some implementations over-replay. Others under-replay. Both create risk.

Over-replay problems include:

  • Resending messages outside the requested range.
  • Replaying application messages without proper duplicate signaling.
  • Replaying administrative traffic in a way the peer does not expect.

Under-replay problems include:

  • Omitting a stored message from the requested range.
  • Sending a gap fill for content that should be retransmitted.
  • Advancing sequence numbers before recovery is complete.

There is also the issue of ordering under load. A session engine with poor serialization around storage and send paths may generate traces that look valid in isolation but fail when multiple events overlap. Heartbeats, application flow, resend handling, and disconnect timers all compete for ordering. If the state machine is loose, one recovery path masks another bug.

Another common mistake is treating the protocol transcript as secondary to the business outcome. If an order appears to make it through, teams sometimes ignore session anomalies. That is short-term thinking. In production, the transcript is the evidence you rely on when counterparties disagree, regulators ask for message flow, or internal teams debug a break. Session integrity matters even when business flow seems fine.

What this demonstrates about system design

FIX Lab demonstrates a design principle that applies well beyond trading protocols. Recovery logic should be observable, deterministic, and testable as a first-class feature.

Observable means you can see the gap, the request, the replay, and the restored state. Hidden recovery logic invites false confidence.

Deterministic means the same fault produces the same protocol behavior. If recovery depends on timing accidents, your integration is fragile.

Testable means you can induce a fault on purpose. Systems built only for successful flow often fail the first time the network, storage, or peer behaves imperfectly.

This matters because message-oriented systems often claim reliability while exposing little evidence. The right standard is narrower. Show the failure. Show the state transition. Show the convergence. A lab centered on one lost FIX message teaches this clearly.

It also teaches restraint. You do not need a large scenario to learn a lot. A single dropped message is enough to inspect the quality of sequence management, resend behavior, replay handling, and final consistency. If those fundamentals are weak, scaling up only hides the problem.

What to watch next

The next step is breadth. Try different loss points, reconnect boundaries, and resend ranges. Watch whether the same recovery discipline holds.

Then watch the handoff from session to application state. A session can recover while downstream processing still mishandles replayed messages. That boundary deserves the same level of inspection.

If a tool like this changes how you test, it has done its job. You stop asking whether messages flow on a good day. You start checking how the engine behaves when the sequence does not.