Instrumentation
Across applications, infrastructure and logs, starting from the handful of signals that show whether the service is genuinely healthy from the outside.
Technology partners
Interkey instruments the estate, builds the dashboards and alerting teams will act on, controls what log volume does to the bill, and operates the platform afterwards — with the residency question settled before the design is fixed.
Implementation team in Riyadh, working in Arabic and English.
In short
Interkey assesses, implements and operates Datadog for organisations in Saudi Arabia. An engagement covers instrumenting applications, infrastructure and logs; designing the dashboards and alerts an operations team will act on; settling where telemetry lands before the design is fixed; and running or handing over the platform afterwards. Implementation team in Riyadh, working in Arabic and English.
It is addressed to infrastructure and platform leaders, SRE and operations teams, cloud and application owners, and the technology decision-makers funding the rollout. Interkey is the implementer here, not the vendor: Datadog supplies the platform and its own documentation, and this page covers the work of making it fit an estate that already exists. If you are weighing platforms rather than rolling one out, the practice page starts a step earlier.
The work
The recognisable failure is a deployment where agents are everywhere, hundreds of dashboards exist, alerts fire constantly, and the team still finds out from a user. The tool was not the problem; nobody decided what it should watch. That decision is where this work starts, and it runs in this order.
Across applications, infrastructure and logs, starting from the handful of signals that show whether the service is genuinely healthy from the outside.
Built for the people who will actually look at them, which usually means fewer and more specific.
Designed around consequence rather than threshold, so an alert firing at 3am means someone genuinely needs to be awake.
Agent versions, tagging discipline, cost review and alert tuning after go-live, because an unowned observability platform decays in months.
On alert fatigue.
Alert design
Every alert is a standing claim on somebody’s night, so each one has to pass the same four tests before it ships. Deployments nobody trusts have usually failed the second and third.
| The test | What it means in practice |
|---|---|
| Does it mean something is wrong? | A threshold being crossed is not the same as the service failing. Alert on symptoms a user would notice — error rate, latency, queue depth against capacity — rather than on every resource metric that happens to have a number. |
| Is there something to do about it? | If the answer at 3am is to acknowledge it and go back to sleep, it is not an alert. It is a dashboard, a weekly report, or nothing. This is the test that removes the most noise. |
| Does the receiver know what to do? | The alert names the affected service, the likely cause and the first thing to check, or it links to something that does. An alert that starts an investigation from zero costs more than it saves. |
| Will it still be true in six months? | Thresholds set against last quarter’s traffic fire constantly after growth; alerts pinned to a specific hostname break at the next migration. Both are why alert quality is reviewed as ongoing work rather than set once at go-live. |
Delivery method
Phases rather than week numbers: the honest duration depends on the size of the estate and how much your own team takes on. What does not vary is the order.
Phase 01
Inventory what runs and which services carry the business, and settle the residency question before any agent ships a byte: what may be collected, what must be filtered, and where it is allowed to land.
Phase 02
Agents and integrations across the first tranche of hosts, clusters and applications, with the network path agreed with the security team rather than worked around.
Phase 03
Service-level dashboards for the teams that own the services, an alert set designed around consequence, and escalation wired into however the organisation actually responds: on-call rotas, ticketing, messaging.
Phase 04
The period that decides whether the platform survives: alert tuning against real incident history, log pipelines reviewed so the bill tracks value, and a handover that leaves your team able to run it.
Before design is fixed
These are settled at design time, not discovered in month six. Each one changes the shape of the rollout rather than its size.
| Decision | What forces it | Where it lands |
|---|---|---|
| Residency | Datadog is SaaS, operated from published regions, and as of August 2026 none is in the Middle East. | Decide deliberately which telemetry leaves the Kingdom, which is filtered or redacted first, and which stays behind entirely. Where that rules a SaaS platform out, it is said plainly. |
| Hybrid estate | Long-lived systems in a local data centre alongside cloud and Kubernetes workloads. | The awkward half is the half quick-start guides skip. It is where an implementation partner earns its keep. |
| Log volume | Per-host pricing is predictable; ingested log volume is not, and a few chatty systems shipping debug noise decide the bill. | Decide which logs must be searchable, which only need to exist, and which never leave your infrastructure — then build the pipelines to enforce it. Cost design belongs in implementation, not a rescue project. What drives the bill, in detail. |
| Portability | OpenTelemetry instrumentation outlives any single vendor decision, and Datadog ingests it. | Instrument portably where it is cheap, use the vendor agent where it earns its keep, and write down which is which. |
| Day-two ownership | The quiet failure is ownership: dashboards age, agents drift, new services ship uninstrumented. | Somebody owns alert quality, cost review and instrumentation standards as ongoing work. Interkey operates it from Riyadh Sunday to Thursday, or backs the team that does. |
Before you ask for a quote
Three things set the size of a Datadog rollout, and none of them is the licence count. The first is how much of the estate is uniform: a fleet of similar Kubernetes workloads instruments once and repeats, while twenty long-lived systems each configured differently are twenty small projects. The second is how much your own team will take on — a platform team that instruments its own services changes the shape of the engagement entirely. The third is whether residency is a live constraint, because that is settled before design rather than during it.
None of that needs a meeting to establish. A rough estate inventory, an honest answer on who will own the platform afterwards, and a note of which data classes cannot leave the Kingdom are enough to scope a first conversation properly — and enough to tell whether the answer should be Datadog at all. The implementation questions buyers actually ask cover duration, cost drivers and hybrid coverage in more detail.
Buyer questions
Direct answers to the questions that come up in real evaluations. Anything missing, ask us at the bottom of the page.
Interkey does the assessment, the instrumentation, the dashboard and alert design, and the cost controls, then either operates the platform or hands it over with documentation. What stays with your team is the decision-making: which services matter, what an alert should mean, what your residency position is, and who carries the pager. Those cannot be outsourced without the result drifting away from how the business actually runs.
Yes, and usually it shortens it. The work starts from what is already instrumented rather than from nothing, and the first phase becomes an audit: what is being collected, what it costs, which alerts fire and whether anyone acts on them. Existing dashboards are kept where they earn their place. A deployment that is already in production is a starting position, not a reason to begin again.
Interkey implements and operates Datadog for customers in Saudi Arabia, with an implementation team in Riyadh working in Arabic and English. Interkey and Datadog teams both participated in Observability Day in Riyadh on 12 November 2025. Interkey does not claim a certified, authorised or tiered partner status, because that is a statement only the vendor can make — if partner status matters to your procurement process, ask us and we will tell you exactly what we can and cannot evidence.
When the estate is small enough and uniform enough that your own team can instrument it in a fortnight from the vendor documentation, an implementation engagement is hard to justify. When telemetry cannot leave the Kingdom at all under any filtering, the honest answer is a different platform rather than a Datadog rollout — the cases where that applies are set out separately. And when nobody will own the platform afterwards, the rollout is worth deferring until that is resolved.
Next step
Tell us roughly what runs where — cloud, on-premise, clusters — and whether residency is a live question, and the first reply comes from an engineer who has done this here.
Datadog implementation and observability services in KSA
Instrumentation, alert design and day-two operation
Riyadh engineering, Sunday to Thursday
Arabic and English, in the network’s own time zone
Observability Day, Riyadh 2025
Interkey and Datadog teams participated