Instrumentation
Across applications, infrastructure and logs, starting from the handful of signals that show whether the service is genuinely healthy from the outside.
Technology partners
Interkey instruments the estate, builds the dashboards and alerting teams will act on, controls what log volume does to the bill, and operates the platform afterwards — with the residency question settled before the design is fixed.
The work
The recognisable failure is a deployment where agents are everywhere, hundreds of dashboards exist, alerts fire constantly, and the team still finds out from a user. The tool was not the problem; nobody decided what it should watch. That decision is where this work starts, and it runs in this order.
Across applications, infrastructure and logs, starting from the handful of signals that show whether the service is genuinely healthy from the outside.
Built for the people who will actually look at them, which usually means fewer and more specific.
Designed around consequence rather than threshold, so an alert firing at 3am means someone genuinely needs to be awake.
Agent versions, tagging discipline, cost review and alert tuning after go-live, because an unowned observability platform decays in months.
On alert fatigue.
Delivery method
Phases rather than week numbers: the honest duration depends on the size of the estate and how much your own team takes on. What does not vary is the order.
Phase 01
Inventory what runs and which services carry the business, and settle the residency question before any agent ships a byte: what may be collected, what must be filtered, and where it is allowed to land.
Phase 02
Agents and integrations across the first tranche of hosts, clusters and applications, with the network path agreed with the security team rather than worked around.
Phase 03
Service-level dashboards for the teams that own the services, an alert set designed around consequence, and escalation wired into however the organisation actually responds: on-call rotas, ticketing, messaging.
Phase 04
The period that decides whether the platform survives: alert tuning against real incident history, log pipelines reviewed so the bill tracks value, and a handover that leaves your team able to run it.
Before design is fixed
These are settled at design time, not discovered in month six. Each one changes the shape of the rollout rather than its size.
| Decision | What forces it | Where it lands |
|---|---|---|
| Residency | Datadog is SaaS, operated from published regions, and as of August 2026 none is in the Middle East. | Decide deliberately which telemetry leaves the Kingdom, which is filtered or redacted first, and which stays behind entirely. Where that rules a SaaS platform out, it is said plainly. |
| Hybrid estate | Long-lived systems in a local data centre alongside cloud and Kubernetes workloads. | The awkward half is the half quick-start guides skip. It is where an implementation partner earns its keep. |
| Log volume | Per-host pricing is predictable; ingested log volume is not, and a few chatty systems shipping debug noise decide the bill. | Decide which logs must be searchable, which only need to exist, and which never leave your infrastructure — then build the pipelines to enforce it. Cost design belongs in implementation, not a rescue project. |
| Portability | OpenTelemetry instrumentation outlives any single vendor decision, and Datadog ingests it. | Instrument portably where it is cheap, use the vendor agent where it earns its keep, and write down which is which. |
| Day-two ownership | The quiet failure is ownership: dashboards age, agents drift, new services ship uninstrumented. | Somebody owns alert quality, cost review and instrumentation standards as ongoing work. Interkey operates it from Riyadh Sunday to Thursday, or backs the team that does. |
Datadog partner in KSA
As named in Datadog’s own partner overview
Riyadh engineering, Sunday to Thursday
Arabic and English, in the network’s own time zone
Observability Day, Riyadh 2025
A joint Datadog and Interkey event
Next step
Tell us roughly what runs where — cloud, on-premise, clusters — and whether residency is a live question, and the first reply comes from an engineer who has done this here.