A dozen client deployments, and the question I opened every one with
Eighteen months doing Splunk Professional Services, with a different client every few weeks. Ticketing SaaS, paper manufacturing, title insurance, rail telematics, county government, healthcare quality — engagements that had almost nothing in common except that I started all of them the same way, and that's the part that made them work.
- Where
- SP6 / Aditum, Splunk Professional Services
- When
- Jan 2020 – Apr 2021
- Role
- Delivery engineer, client-facing
- Stack
- Splunk Enterprise & Cloud, AWS, CIM
The question
Every kickoff call, after the introductions and before the SOW walkthrough, I asked some version of this:
Before we get started, does your team have a picture of what you consider “done”? At what point during our work would you feel the project is at a satisfactory level? I ask because I want to prioritise your team’s success, and agreeing the acceptance criteria up front keeps us pointed in the same direction.
It sounds like a soft question. It isn’t. A statement of work says what will be delivered; it very rarely says what “good” looks like to the people paying for it, and those are different documents. I’ve watched engagements deliver every line item and still leave the client unsatisfied, because what they actually wanted was one thing that never made it onto the list.
I’d usually offer a few candidate answers to get people talking — “we want Splunk healthy enough to put security on it”, “we want a highly available architecture”, “we want to stop having downtime we can’t explain”. Clients almost always picked one and then corrected it into something more specific, which is exactly the point.
The other thing I said early, every time, was how much wouldn’t fit. If the engagement was twelve days and the data-onboarding list was thirty days long, saying so in week one made me the person being straight with them. Saying so in week three would have made me the person who’d overrun.
What the engagements looked like
Tessitura — on-prem to Splunk Cloud, three regions
A net-new Splunk Cloud build replacing on-premises infrastructure spread across a legacy datacentre and AWS, in three geographic regions with separate Cloud environments for each. Eight sites, at minimum one heavy forwarder / deployment server per site.
The work: recreating the on-prem index structure in Cloud, lifting the client’s in-house app and their PCI app across without reconfiguration — their queries had to keep working unchanged, which was a hard requirement and the right one — retention at 90 days before rolling into archival storage, and duplicating data to both old and new environments while the new one was validated. Same dual-write principle I’d later use at scale on the consolidation project.
Migrating into Splunk Cloud also means apps have to pass Splunk’s inspection process before they can be deployed, which is a real scheduling constraint rather than a formality, and worth surfacing to a client before it’s on the critical path.
County of Tooele — replacing a third-party SIEM
A county government moving security monitoring in-house from an outsourced SIEM. Splunk install plus data onboarding, using Security Essentials and InfoSec as a lighter-weight alternative to full Enterprise Security.
Their definition of done was the most interesting one I got, and it wasn’t coverage. They explicitly weren’t worried about getting every endpoint in. What they wanted was to know how to build alerts and use the apps themselves — if we onboarded a server OS and a network OS between us, they’d take it from there. That reframed the whole engagement from a data-onboarding job into a capability-transfer job, and it changed what I spent the hours on.
Babcock & Wilcox — environment review
An industrial energy manufacturer. Engagement readiness review and environment assessment, framed against the CIS controls.
First American — building a data-loss use case from scratch
A title insurance and financial services company wanting to get real value out of their DLP tooling rather than just collecting its alerts. The useful move was to stop thinking about the product and start thinking about the analyst: what would someone actually want to know? Is a person offloading data ahead of leaving? Is an unusual volume moving? Are records leaving that shouldn’t? Those questions became the dashboard — one an analyst could put a username into and get an answer.
Wi-Tronix — sizing an architecture before anyone bought hardware
Rail telematics. They knew they were growing and didn’t know what to build. I worked from their actual license volume and existing kit to a concrete recommendation — a multi-site indexer cluster split across two sites, a search head with a warm standby, and the management tier to go with it — rather than the largest architecture that would fit the budget.
Warren County Telecom — CIS controls, top down
Working through the CIS controls list and mapping each one to something Splunk could actually answer. The first control is inventory of hardware assets, which sounds trivial and isn’t: it meant pulling from Active Directory into a lookup that everything else could join against. Same lesson as the compliance work — a control framework is a detection backlog if you read it that way.
WC Bradley, International Paper, ACA Track, Telligen
The rest of the run, briefly: a net-new Splunk deployment inside a client’s Azure environment with hybrid search bridging old and new during migration; a managed-service engagement where the remit was explicitly keeping Splunk healthy rather than watching alerts; a syslog-ng ingestion build-out with Splunk Security Essentials on top; and an environment health check for a healthcare quality organisation.
Explaining things without dumbing them down
A lot of consulting is translation. Clients are paying for Splunk expertise they don’t have, and the useful move is usually to explain the concept in terms of their data rather than in terms of the product’s vocabulary.
The one I gave most often was the Common Information Model. The version I’d use, more or less verbatim:
The CIM helps you normalise your data to a common standard. So if you have authentication logs in Linux that include a
userfield, but other logs on your Windows server that have auserNAMEfield, you’ll want to make sure both of those have the same name — so Splunk can work more efficiently.
No jargon, one concrete example, and the reason it matters. That explanation got more people to care about CIM mapping than any amount of talking about data model acceleration ever did.
What I took from it
- Acceptance criteria are the deliverable that makes the other deliverables mean something. I’ve used the “what does done look like?” question everywhere since, and I regret every time I’ve skipped it.
- Say what won’t fit, early. Scope honesty in week one is competence; in week three it’s an excuse.
- Sometimes the job is teaching. The county engagement succeeded because I stopped optimising for endpoints onboarded and started optimising for what they could do after I left.
- Breadth compounds. A dozen environments in eighteen months meant I’d seen most of the ways a Splunk deployment goes wrong before I ever had to run one at scale in-house.
What I’d do differently
I’d have written the handover documentation as I went rather than at the end of each engagement. The consulting model rewards leaving a client self-sufficient, and the runbook written on day nine is always thinner than the one that got written on days two through nine. It’s the same lesson the TLS work taught me later, which suggests I should have learned it properly the first time.
Client names appear here because these engagements are ones I can speak to; individual contacts, environment specifics, and anything found during the assessments do not.