Software that survives production.
We build high-performance systems designed for failure, and we find out why yours are slow, hanging or crashing. Senior engineering, from the crash dump to the architecture.
Book a 30-minute callProduction in trouble? Start here
30+ years of engineering experience · Senior engineers on every engagement · Your incidents stay confidential
When production breaks, we find out why.
Hangs under load. Crashes nobody can reproduce. Memory that only goes up. Latency that appears at 10 a.m. and is gone by the time someone looks. These problems live in the running system, and that is where we look, down to the exact thread, lock, allocation or packet responsible.
Live debugging
attachWhen the problem is happening now. We attach to the running process with minimal impact, inspect threads, locks and memory in place, and capture exactly the evidence needed, at the right moment, without a restart that would make the problem disappear.
Post-mortem analysis
.dmpWhen the process is already gone. Crash and hang dumps tell us what every thread was doing at the moment of failure. We read them down to the stack frame, including damaged or incomplete dumps, and the crashes that leave nothing in the logs.
Performance & resource profiling
.etlWhen it works, but too slowly or too expensively. CPU sampling, ETW tracing, allocation and GC analysis, thread pool behaviour, handle and native memory growth: measured in your environment, under your real load.
Networking diagnostics
.pcapngWhen the problem is between machines. Timeouts, connection resets, port exhaustion, DNS and TLS delays, latency that only appears across a firewall or load balancer. We correlate packet captures with what the application was doing at that same millisecond.
WinDbg · PerfView · Windows Performance Analyzer · ETW · dotnet-dump · dotnet-trace · dotnet-counters · Process Monitor · Wireshark
Production Failure & Performance Audit
Fixed scope. Evidence, not opinions.
- Collect — dumps, traces, counters and captures from your environment, planned to keep the impact on production minimal.
- Analyse — down to the code path, the lock and the line responsible.
- Report — root cause with the evidence behind it, ranked fixes, and a plan your team can execute, or that we can execute with you.
Scope and pricing on request, after a 30-minute call.
Book an audit callWhat it usually is.
Production failures repeat. The same handful of root causes hides behind very different symptoms. A few illustrative examples of the patterns we diagnose:
“It freezes every morning at peak, then recovers.”
Live debugging
- Where we look
- a hang dump taken during the freeze; thread pool and lock state.
- What it often is
- thread pool starvation from synchronous waits on asynchronous code. Every worker thread is blocked waiting for work that needs a free thread to complete.
* hang dump of w3wp.exe, taken during the 09:14 freeze
0:000> !threadpool
CPU utilization: 3%
Workers Total: 64
Workers Running: 64
Workers Idle: 0
Work Request in Queue: 1287
0:000> ~*e !clrstack
OS Thread Id: 0x2c48 (41)
System.Threading.Tasks.Task.Wait()
Orders.PricingClient.GetQuote()
Orders.QuoteController.Get(Int32)
* 64 of 64 workers blocked in Task.Wait(); none left to finish the work
“It crashes once a week. Nothing in the logs.”
Post-mortem
- Where we look
- a crash dump captured automatically on the next occurrence.
- What it often is
- a stack overflow or a fault in native code. Both kill the process before any logging code gets to run.
* crash dump written by WER LocalDumps on the next occurrence
0:017> .exr -1
ExceptionAddress: 00007ffd3a1b2c4e (rules!Node::Evaluate+0x1e)
ExceptionCode: c00000fd (Stack overflow)
ExceptionFlags: 00000001
0:017> kn 3
# Child-SP RetAddr Call Site
00 0000002b`f2a03fe0 00007ffd`3a1b2d17 rules!Node::Evaluate+0x1e
01 0000002b`f2a04050 00007ffd`3a1b2d17 rules!Node::Evaluate+0xe7
02 0000002b`f2a040c0 00007ffd`3a1b2d17 rules!Node::Evaluate+0xe7
* the same frame, thousands deep: unbounded recursion
“Memory climbs for days until the service is restarted.”
Profiling
- Where we look
- two or three dumps hours apart, compared by type and by retention path.
- What it often is
- an event handler or cache that keeps every object alive, a blocked finalizer thread, or large object heap fragmentation.
* third of three dumps, six hours apart
0:000> !dumpheap -stat
MT Count TotalSize Class Name
00007ffa8c1d2f10 2108733 168698640 Orders.OrderSnapshot
0:000> !gcroot 000001d4a8c13e40
-> 000001d4a2a01020 System.Object[]
-> 000001d4a8c0f2b8 Pricing.PriceFeed
-> 000001d4a8c0f330 System.EventHandler<PriceChangedEventArgs>
-> 000001d4a8c13e40 Orders.OrderSnapshot
0:000> !finalizequeue
Ready for finalization 0 objects
* every snapshot stays subscribed to a static PriceFeed event
“Requests to the partner API start failing under load.”
Networking
- Where we look
- connection tables, TCP states and a packet capture at the moment of failure.
- What it often is
- ephemeral port exhaustion, from opening a new connection per request instead of reusing a pool.
# on the app server, during the failure window
PS> Get-NetTCPConnection -State TimeWait | Measure-Object
Count : 16231
PS> netsh int ipv4 show dynamicport tcp
Protocol tcp Dynamic Port Range
---------------------------------
Start Port : 49152
Number of Ports : 16384
# 16,231 of 16,384 ephemeral ports parked in TIME_WAIT
“The first request after a deploy takes 30 seconds.”
Networking
- Where we look
- an ETW trace of the startup and the outbound connections it makes.
- What it often is
- certificate revocation checks timing out against an endpoint the server cannot reach, multiplied by every TLS connection made at startup.
# CAPI2 logging is off by default: switch it on first
PS> wevtutil sl Microsoft-Windows-CAPI2/Operational /e:true
# recycle the app pool, send one request, then read the log
PS> wevtutil qe Microsoft-Windows-CAPI2/Operational /c:1 /rd:true /f:xml
<EventID>53</EventID> … <Level>2</Level>
<URL scheme="http">http://crl.pki.example.com/issuing-ca.crl</URL>
<Result value="800705B4">This operation returned because
the timeout period expired.</Result>
# CRL endpoint unreachable: every chain build waits out the timeout
“CPU is at 100% and nobody changed anything.”
Profiling
- Where we look
- CPU sampling with full stacks, then the data that drives the hot path.
- What it often is
- an algorithm that was fine at last year's data volume, or a retry loop that turned a downstream slowdown into a storm.
# while the CPU is pinned
PS> wpr -start CPU
PS> wpr -stop cpu.etl
# cpu.etl in WPA: CPU Usage (Sampled), by process and stack
% Weight Stack
71.36 |- Billing.dll!Billing.Reconciler.Match
70.92 | |- System.Linq.Enumerable.Contains<Int64>
# Contains() inside a loop: O(n²) on four times last year's rows
Recognise one of these? Tell us what you're seeing
Built to survive failure.
Every system fails eventually: a dependency times out, a disk fills, a network blips, traffic triples. The question is what happens next. We design for that moment first, then make the system fast.
Resilience
- Timeouts, retry budgets and circuit breakers that contain a failure instead of spreading it.
- Bulkheads and backpressure, so one slow dependency cannot take the whole system down.
- Idempotent operations and safe recovery, so a retry never charges a customer twice.
- Graceful degradation: the system keeps doing the important work when parts of it are down.
- Failure testing: we break it on purpose before production does.
Performance
- High-concurrency services with predictable latency under load.
- Allocation-aware, GC-friendly .NET code on the hot paths.
- Latency budgets set per request, measured, and enforced.
- Load testing to the breaking point, so you know where it is.
What we deliver
- New systems — high-performance services and platforms, designed for failure from day one.
- Hardening — taking an existing system through a failure review and fixing what it finds.
- Modernisation — moving legacy .NET Framework and Windows workloads to modern .NET without losing what already works.
- Integration — APIs, middleware and automation between systems that were never designed to talk to each other, Microsoft Power Platform included.
AI that holds up in production.
Most AI projects work in the demo. Production is different: model calls that take 20 seconds, rate limits at peak hour, costs that scale faster than usage, answers that change from one run to the next, and data that is not allowed to leave the building. These are performance and resilience problems, and that is our work.
What we build
- LLM integration into existing systems — with timeouts, fallbacks, caching and cost control from the first release.
- Evaluation harnesses — automated tests that tell you whether a prompt or model change made things better or worse.
- Observability — latency, cost and quality per request, so AI behaviour is measured.
- Local and private models — when the data has to stay on your infrastructure.
- AI-assisted operations — automation that triages, summarises and routes work, with a person approving what matters.
How we work.
- Senior engineers on every engagement. The person who scopes the work is the person who does it. No sales layer, no hand-off.
- Evidence first. Every conclusion comes with the dump, trace or measurement behind it.
- Confidential by default. We sign your NDA before we see anything, and we never publish client work: no logos, no case studies, no names. Your incidents stay yours.
- Fixed scope to start. The audit is the low-risk way to find out whether we are the right fit.
- We work with your team. Your engineers see the method and keep the knowledge.
- Remote across Europe, on-site in the Lisbon area when it helps.
Tell us what's failing.
A short description is enough. We reply within one business day.
- geral@ncode.pt
- Address
- Mafra, Portugal