正在加载内容...

963963 Chat Review Portal Independent coverage of news

Common Mistakes When Evaluating Monitoring Alerts

By Robert Hayes · · 1207 words
Common Mistakes When Evaluating Monitoring Alerts

For edge caching, the constraint matters more than the feature list. The first thing to settle is the failure mode, not the happy path. Teams working on edge caching usually discover this the hard way. Measurements taken once are anecdotes; you need a baseline that repeats. Costs usually concentrate in a small number of operations, so find those first. This is most visible in edge caching.

The interesting number is not the average, it is the 99th percentile. The same reasoning holds for load balancing. For load balancing, the constraint matters more than the feature list. Adding a cache in front of a slow query is a fix; fixing the query is a cure. Teams working on load balancing usually discover this the hard way. Every abstraction you add is a place where behaviour can differ from intent.

Configurations should be reviewable in a diff, not only in a console. This is most visible in schema migration. Consider schema migration specifically. The best time to add an index is before the table gets large. Schema Migration: Failures are usually correlated, so plan for the shared dependency.

Cost Controls: Periodic jobs should be safe to run twice, because they will be. Cost Controls: You rarely need a new component to fix a boundary problem. Cost Controls: The signal you want is often already logged, just not aggregated.

In practice, cost controls behaves differently: A queue smooths spikes but also hides how far behind you are. Retries without jitter turn a small outage into a large one. The same reasoning holds for cost controls. For cost controls, the constraint matters more than the feature list. Separating the reads from the writes buys room to change either side.

Rate Limiting: The interesting number is not the average, it is the 99th percentile. Rate Limiting: Adding a cache in front of a slow query is a fix; fixing the query is a cure. Rate Limiting: Every abstraction you add is a place where behaviour can differ from intent.

You can often replace a coordination problem with an idempotency key. The same reasoning holds for search indexing. For search indexing, the constraint matters more than the feature list. Anything that grows without a bound will eventually hit one. Teams working on search indexing usually discover this the hard way. Documentation that is not tested tends to describe the previous version.

Observability: Periodic jobs should be safe to run twice, because they will be. Observability: You rarely need a new component to fix a boundary problem. Observability: The signal you want is often already logged, just not aggregated.

Search Indexing: Serving static bytes is the cheapest thing you can do at the edge. Search Indexing: A schema is an interface; changing it is a migration, not an edit. Search Indexing: Track the denominator as carefully as the numerator.

Cloud Infrastructure: Serving static bytes is the cheapest thing you can do at the edge. Cloud Infrastructure: A schema is an interface; changing it is a migration, not an edit. Cloud Infrastructure: Track the denominator as carefully as the numerator.

Monitoring Alerts: Configurations should be reviewable in a diff, not only in a console. Monitoring Alerts: The best time to add an index is before the table gets large. Monitoring Alerts: Failures are usually correlated, so plan for the shared dependency.

Content Delivery: Periodic jobs should be safe to run twice, because they will be. Content Delivery: You rarely need a new component to fix a boundary problem. Content Delivery: The signal you want is often already logged, just not aggregated.

Teams working on log analysis usually discover this the hard way. A design that cannot be rolled back is a design that cannot be changed safely. Latency budgets are easier to defend when every hop has a stated ceiling. This is most visible in log analysis. Consider log analysis specifically. Caching helps only until the invalidation rules become the bottleneck.

Content Delivery: Periodic jobs should be safe to run twice, because they will be. You rarely need a new component to fix a boundary problem. That applies to content delivery as well. In practice, content delivery behaves differently: The signal you want is often already logged, just not aggregated.

Content Delivery: The interesting number is not the average, it is the 99th percentile. Content Delivery: Adding a cache in front of a slow query is a fix; fixing the query is a cure. Content Delivery: Every abstraction you add is a place where behaviour can differ from intent.

Use direct, ordinary language. For example, ask, “Would you like to continue?” or “Are you comfortable with this?” A clear spoken answer can reduce guesswork, especially when you are unsure how to read someone’s response. Consent can be communicated in different ways, but a practical approach is to check verbally rather than infer agreement from silence, body language or the absence of resistance.

Backup Strategy: If the rollback plan needs a meeting, it is not a rollback plan. Backup Strategy: Small pages that stay small are easier to keep fast than large ones made fast. Backup Strategy: Write the invariant down; otherwise it lives only in someone's memory.

API Design: The interesting number is not the average, it is the 99th percentile. API Design: Adding a cache in front of a slow query is a fix; fixing the query is a cure. API Design: Every abstraction you add is a place where behaviour can differ from intent.

Crawl Budget: Configurations should be reviewable in a diff, not only in a console. Crawl Budget: The best time to add an index is before the table gets large. Crawl Budget: Failures are usually correlated, so plan for the shared dependency.

Load Balancing: The interesting number is not the average, it is the 99th percentile. Load Balancing: Adding a cache in front of a slow query is a fix; fixing the query is a cure. Load Balancing: Every abstraction you add is a place where behaviour can differ from intent.

For data pipelines, the constraint matters more than the feature list. If a metric has no owner, it will drift until it causes an incident. Teams working on data pipelines usually discover this the hard way. The cheapest optimisation is usually removing work nobody asked for. Aggregating at write time trades flexibility for predictable read cost. This is most visible in data pipelines.

Consider backup strategy specifically. You can often replace a coordination problem with an idempotency key. Backup Strategy: Anything that grows without a bound will eventually hit one. Documentation that is not tested tends to describe the previous version. That applies to backup strategy as well.

Consider storage tiers specifically. You can often replace a coordination problem with an idempotency key. Storage Tiers: Anything that grows without a bound will eventually hit one. Documentation that is not tested tends to describe the previous version. That applies to storage tiers as well.

Teams working on search indexing usually discover this the hard way. Serving static bytes is the cheapest thing you can do at the edge. A schema is an interface; changing it is a migration, not an edit. This is most visible in search indexing. Consider search indexing specifically. Track the denominator as carefully as the numerator.

Related reading