The “Do No Harm” Analytics Manifesto
It's been 4 years and a day since my last blog post (a.k.a. way too damn long), and I apologize for that! But a lot has been going on since I last posted; I've switched jobs a few times, become a father of twins, spoken at Adobe Summit, and I have even shaved my head!
I'm hoping, however, that this blog post will make the wait worth it, as it's been something I've been personally noodling on since my time at Kroger. And it is all thanks to a phrase someone there coined (I honestly can't remember who) that stuck with me: as part of a good analytics practice, we should "Do No Harm" to our existing analytics, or others' analytics / data.
This blog is the result of me trying to take a stab at creating a blueprint of sorts for what the hell this even means for you and your organization, your client's organization, or any organization that you are wanting to implement a successful data-driven analytics practice at. I've broken this out into "Core Beliefs", the commandments that you should follow with this ideology and the three "Pillars" on which the strong foundation for your organization's analytics practice will be able to build and scale without worrying that it will all come down like a Jenga tower, and help support the core beliefs of this ideology.

Photographic proof
Core Beliefs / Commandments
1. Relevance is Perspective
i.e., "Just because you're not using it, doesn't mean it's not important"
If I had a nickel for every time I heard "Well I only care about 'x' metric, my feature captures that, why can't we go live?" from a product manager, I'd have a pretty good start on my kids' college fund. In the process of rushing out the door, and getting the feature working ASAP, yes, the metrics you care about are firing as requested and can be reported on, but it is now fully breaking other features / metrics from the last release. You can take this and flip it the other way as well, as I've had analysts tell me "Well, let's just rip out this eVar so we can repurpose it for 'y', because I've never used it for my reporting."
At the end of the day, there was an end consumer for the data at some point, or stakeholder who requested it (even if they never used it). It's important to understand the context & rationale for the metric / dimension being captured. It's also important that we consider that a measurement plan can often span multiple products & teams, and it's the responsibility of all stakeholders to ensure that data collection is consistent throughout.
2. Newton's Third Law Also Applies to Analytics
i.e., "Every Change Warrants an Equal and Opposite Reaction"
This is going to sound obvious, but in any field where we find complex & innovative ways to parse & process data, we sometimes can miss the forest for the trees, so it's worth stating very clearly.
If you move a link to a higher spot on a page or menu than another link, yes you will raise visibility & likely clicks for that link. However, you are also going to lower visibility and likely clicks for the link that ends up getting pushed down for that link to rise up.
In larger organizations, you can have situations where two different teams manage the end services of these two links, and while the first team who had the link raised in the navigation is ecstatic to see the increased volume to their service, the second team who is now seeing a consistent decline as this feature is tested and rolled out may be out of the loop on this change, and only find out months later after it's been made and rolled out.
Obviously, this is inherent in any change that is made to some extent, but it reinforces that we must make decisions with the end consumers in mind; if they can't find what they need within your walled garden (that's relevant and useful to them regardless of where they are at in the funnel), they will go elsewhere. Resist the urge to operate in silos when making a change and to focus only on the positive impact of a change.
3. Analytics is a Garden
Prune & maintain with care, don't pull up items that aren't "weeds"
Pruning is essential for health and maintenance of your analytics architecture. Deprecated metrics, dimensions, eVars, props, goals, etc. are dead plants that block space for new growth, or get mistaken for actual plants.
Metrics captured incorrectly or defects / issues in your implementation will grow like weeds choking the life out of valuable insights and hurt your data confidence, and they need to be eradicated like the sharp, painful thistle they are.
However, regardless of the state of your current analytics practice, it is essential to resist the urge to rip out EVERYTHING (or get the industrial grade weed killer if we're still using the garden metaphor) and only move forward with this as a last resort. Bare soil will take much longer to sprout life than seedlings / small plants.
4. If You Capture Everything, You Capture Nothing
I'm pretty sure this is what Sun Tzu would say if he worked in data
As your organization grows its analytics maturity, it becomes very tempting to fall into the trap of "Why aren't we capturing everything?". The problem with this ideology is that your definition of "everything" will be different from probably every other stakeholder's definition of "everything". It's especially tempting to fall into this trap in the "Age of AI" where you're one Claude prompt away from "you should be collecting ALL clickstream data, why aren't you?".
I dive into this more in my blog post "The Case Against Clicks" but a great example is if we are to track a form, you could track interactions on form fields, or you could track EVERY SINGLE KEYPRESS. Technically, the former isn't tracking 'everything', but one would argue that knowing if a user typed 3 letters vs. 5 letters in a first name field is a pretty useless key-value pair for most scenarios that senior leadership will ask regarding success of the form.
I'd also argue that while there are more valuable insights derived with more data collected, that there is a finite point where we reach diminishing returns, or the amount of data collected that isn't relevant to answering business questions becomes so great that it actually becomes harmful to the number of insights we can derive. And this is before we even jump into the expense of server calls!
5. Anything Worth Measuring is Worth Measuring Right
How to Prevent "Fear and Loathing" in Your Analytics
If you want to measure something, it's worth delaying a release to properly measure success rather than launching with missing analytics & gaps in important measures of success. I'm obviously a biased source, but I am willing to bet that you can't think of a situation where that occurred and you were able to course-correct / pivot effectively based on limited data (or encourage other stakeholders to adjust their strategy).
Too often in the age of rapid release cycles, fast approaching deadlines, and iterative releases, it's easy to overlook non-functional issues like analytics when you see a console error that breaks the whole application submission process. But I'd argue that the impact to data confidence and insights, especially downstream, isn't worth it. Analytics should be part of your MVP.
The Pillars
Governance
Know your data's journey; where you're going, where you've been, and what wrong turns you have made. Enable trust in the sources so you don't waste time arguing the validity / accuracy of your data.
I've recently become familiar with a concept called "Shift-Left Governance" and the last time I had a concept resonate so profoundly with me was when I first heard about the concept of Event-Driven Data Layers (which, if this is your first experience with my content, I consider to be the inflection point of when I really GOT analytics data capture), so if you are not familiar with this philosophy and mindset, I HIGHLY encourage you to learn more about it.
The one-line summary of this philosophy highly simplified is that we "collect data as it should be as close to the ingestion source as possible". While it didn't have this name at Kroger (at least to my knowledge), it was a tenet that was enforced by having desk checks / analytics QA by our team whenever developers were releasing a feature to ensure analytics was firing correctly, and minimizing analytics hot-fixes or overrides whenever possible. I also want to be clear, it didn't mean that everything was coming up roses for us in that regard, but that we had a North Star guiding how we should handle our validation and at what stage in the process, and the team implementing the analytics collection based on their events was on the ground level validating every feature release. It also means ensuring your development / engineering teams are onboard with validation of their analytics and understand the importance of the structure of it, even if it should never be their priority or focus compared to functional defects.

This is not entirely on DevOps or Engineering teams taking on this task without support or having to hunt down someone on the analytics team to QA; as implementation professionals or analytics / martech engineers, we need to make this as simple and available as possible to implement, by creating structured, grounded rules and documentation on them that can be applied whenever possible to minimize the need for manual QA efforts. Designing data layers, ensuring uniqueness of variable keys, and managing the upstream mapping so engineering teams can implement against an event contract with confidence that the data will be tracked as expected is crucial to this workflow.
Relevance
Does your data provoke new thought processes / give greater visibility and understanding of the customer experience or just reinforce biases?
I've heard "real-time" and "all-encompassing" so many times in the age of AI that I fear I may have developed a knee-jerk reaction to want to slap someone when they utter these words. I joke (somewhat...), but "relevance" doesn't mean that it has to be "the most up-to-date data point we have", although for some instances it is nice / required. It just means that we have confidence that the data is not invalid (i.e., user experience sentiment about a much-maligned UX / UI flow that was removed 4 releases ago for example, should not be used for latest updates if we have not polled user sentiment since trying to address the changes).
For the "all-encompassing" part, please see core belief #4. If you feel that there is valid data that is being missed that is needed to determine proper relevance for business success, a better question is "what don't we know today?" to start with, as you may find out you actually do know the answer based on existing data, or you will have concrete data points identified to ingest for filling this gap.
The last piece I'd stress here is that data should not be used to reaffirm existing beliefs / business strategy, but to move the actionable levers that you can control on items you DON'T know the best route forward. Data that reinforces or merely confirms already held beliefs, validates the already validated, or points towards actions that cannot be taken is unactionable and therefore worthless. A real-life example here:
- There was a legal / compliance requirement to add to an existing form a checkbox that loads a full page overlay of legalese that had to be accepted rather than just silently "accepted"
- There was a request that came to my team to add "analytics tracking for clicks on the checkbox, and clicks inside the form as they scrolled" as they wanted to prove that this additional compliance step would hurt form interaction and conversion
- When we asked if the goal was to identify an alternative way to display this legalese / how they wanted to show the metrics in the measurement plan, we heard the response, "Oh, no this cannot be done any other way. Legal has red-lined it, even if we proved it dropped form conversion to 0%, they won't budge."
Collecting this data is completely irrelevant if that is a compliance / legal requirement; you aren't going to be able to move the lever (modify the display of the legalese) because you've been told that is not possible, and a logical inference is that we aren't going to see form submissions "increase" due to adding an extra step to the conversion flow. You will know what date this feature rolls out and you can monitor existing analytics collection to confirm understanding without collecting anything new. In addition, collecting data to reaffirm biases / expected behavior sets the expectation that you need to do this to confirm other expected behavior, which dilutes the value and ROI on the data you collect as an organization. You should not aim to reinvent the wheel with each release or implementation.
Audience
Know your end consumers and their roles, functions, and where they slot into the organization
One of the core beliefs of this mindset is "Relevance is Perspective", and even though we've already hit on one of the pillars being relevance itself, audience makes up the other half of this core belief, making the most direct connection between core beliefs and pillars in this manifesto.
Your engineering teams will have a different view and understanding of data from your product teams, and the importance of certain data points based on ability to impact plays a role in this. Showing legal page load speed in reporting is likely not giving them any actionable insights that they can influence; bogging down executive dashboards with detailed homepage interactions that result in the whole Analysis Workspace taking 5 minutes to fully load cuts the impact and messaging of the important "how are we doing, at a high-level?" for continued monitoring and ensuring the right trajectory for initiatives is maintained.
Audience also goes hand-in-hand with data storytelling, and how you should build your reporting and dashboards, and more importantly for the people building the data collection pipelines, how you collect these before the feature even releases, so you aren't left flying blind. Too often the audience is such a late afterthought in distributed organizations that by the time you find out who they are, it's too late to address them correctly or fulfill their needs when building reports.
Hitting at other core beliefs, oftentimes a feature that one team is updating can also cause ripple effects downstream that impact other organizational team sources of data, but because they are not the audience for those metrics, they might completely miss this via their organizational silo, which in my opinion is quite possibly the supervillain in this manifesto.
Silos
The saboteur or antithesis to this entire mindset; the vinegar to the baking soda, the water in the fuel line, and the arterial plaque in the heart of your data organization
Because I can't stress it hard enough: SILOS WILL CAUSE YOUR FOUNDATION TO CRUMBLE. They are the earthquake that can smash any (or all) of the three pillars and bring the whole damn building down.
They result in stakeholders missing the audience (or potentially ignoring them entirely), sneaking through governance and data quality processes via tribal knowledge and "oh we just do it this way" mindset, and questioning relevance by not working from a shared organizational understanding. They were damaging enough pre-AI but they are now cataclysmic; AI is an accelerator, and that can be both positive and horribly negative depending on how it is used. Asking Claude to derive reporting insight in silos when data is not accurately labeled or validated, or worse duplicated / deprecated, can result in valid data that is not known being missed, and alternative solutions and pipelines architected that previously wouldn't have been possible, and then you're contending with determining which weed-like plant is the true crop in our analytics garden.
One of my favorite quotes I learned while at Razorfish / Publicis Groupe seems the best way to end this section:
No Silo, No Solo, No Bozo
- Maurice Lévy
Behold the "Do No Harm" Diagram
Where Do I Start?
If you're still reading this, then perhaps we are kindred souls and this post is not just the ravings of a madman into the void. You might also be wondering, "Cool, this seems like an enormously vague ideal dream world. How do we even attempt to push this forward?"
The good and bad news I have for you is that it's doable, and anyone who has a touchpoint in the data process can assist, but it all starts with data lineage & filling the gaps in your documentation, and identifying what you don't know about the important data that is collected in your organization (you individually or you collectively at your whole organization), so that you may attempt to discover how you will come to "know".
That's the fancy, philosophical way of putting it. The truth is less glamorous: pulling rows of data from analytics systems, identifying various permutations of values and creating enumerations from them, etc. But the good news is that you don't have to do this globally; start with the most important metrics & dimensions and work your way down in importance; ideally you'll uncover events that you can validate and adjust as groups, especially if they're all coming from the same source.
A great example that is relevant to many, and what steps could look like:
Marketing Channels & UTM Parameters
- Identify Existing UTMs Passed / Deployed
- You need to determine what is passed / currently active as a campaign so that regardless of final structure, you've added support for the current state as best possible (with the exception of active conflicts / duplications)
- Identify Downstream Stakeholders / Unify Artifacts
- This one is where it gets tricky; who is using each UTM or UTM classification?
- Do all teams have the same agreement / understanding?
- If not, where do they collide?
- When there is a collision, what is the best way of remediation?
- Frankly, this is where it most often falls apart, but you must persevere. Find the places for compromise when there are disagreements, find logical rules and patterns to use as a base for building a foundation, and work most important campaigns / highest volume ones in first
- This one is where it gets tricky; who is using each UTM or UTM classification?
- Make it simple / future proofing
- Prevent the recurrence of the previous state
- It may be that no one ever did this exercise to begin with, and now that it's been done, you'll never encounter this problem again. In the likely situation that is not the case, address the root problems that allowed this to branch out / future-proof for them.
- Create as much documentation as you can actively maintain
- You need to make it easy for people to know the new world order for these values, but it also needs to be maintainable, as documentation that is not updated regularly will slide back into the past state.
- People and processes often follow the path of least resistance; do whatever is in your power to make compliance not only beneficial, but to make clear impacts if not followed (broken reporting, data outages, etc.)
- Prevent the recurrence of the previous state
Final Thoughts
None of this will happen overnight, but every dimension you document, every silo you rip open, and every release you keep from breaking someone else's data moves you one step closer to an analytics practice that truly does no harm. With your analytics garden, you can start small, prune with care, and remember: the goal isn't the world's most beautiful garden, just one that doesn't get wrecked every time someone new picks up a shovel.
I'd love to hear if others have similar approaches or mindsets for their organizations, so please feel free to let me know if I'm completely out of my mind or if this resonates with you. And I promise that the next post won't take another four years.
