📦 deps(thirdparty): update snapshots
This commit is contained in:
+458
@@ -0,0 +1,458 @@
|
||||
---
|
||||
name: distributed-tracing
|
||||
description: "Implement distributed tracing with Jaeger and Tempo for request flow visibility across microservices."
|
||||
risk: critical
|
||||
source: community
|
||||
date_added: "2026-02-27"
|
||||
---
|
||||
|
||||
# Distributed Tracing
|
||||
|
||||
Implement distributed tracing with Jaeger and Tempo for request flow visibility across microservices.
|
||||
|
||||
## Do not use this skill when
|
||||
|
||||
- The task is unrelated to distributed tracing
|
||||
- You need a different domain or tool outside this scope
|
||||
|
||||
## Instructions
|
||||
|
||||
- Clarify goals, constraints, and required inputs.
|
||||
- Apply relevant best practices and validate outcomes.
|
||||
- Provide actionable steps and verification.
|
||||
- If detailed examples are required, open `resources/implementation-playbook.md`.
|
||||
|
||||
## Purpose
|
||||
|
||||
Track requests across distributed systems to understand latency, dependencies, and failure points.
|
||||
|
||||
## Use this skill when
|
||||
|
||||
- Debug latency issues
|
||||
- Understand service dependencies
|
||||
- Identify bottlenecks
|
||||
- Trace error propagation
|
||||
- Analyze request paths
|
||||
|
||||
## Distributed Tracing Concepts
|
||||
|
||||
### Trace Structure
|
||||
```
|
||||
Trace (Request ID: abc123)
|
||||
↓
|
||||
Span (frontend) [100ms]
|
||||
↓
|
||||
Span (api-gateway) [80ms]
|
||||
├→ Span (auth-service) [10ms]
|
||||
└→ Span (user-service) [60ms]
|
||||
└→ Span (database) [40ms]
|
||||
```
|
||||
|
||||
### Key Components
|
||||
- **Trace** - End-to-end request journey
|
||||
- **Span** - Single operation within a trace
|
||||
- **Context** - Metadata propagated between services
|
||||
- **Tags** - Key-value pairs for filtering
|
||||
- **Logs** - Timestamped events within a span
|
||||
|
||||
## Jaeger Setup
|
||||
|
||||
### Kubernetes Deployment
|
||||
|
||||
```bash
|
||||
# Deploy Jaeger Operator
|
||||
kubectl create namespace observability
|
||||
kubectl create -f https://github.com/jaegertracing/jaeger-operator/releases/download/v1.51.0/jaeger-operator.yaml -n observability
|
||||
|
||||
# Deploy Jaeger instance
|
||||
kubectl apply -f - <<EOF
|
||||
apiVersion: jaegertracing.io/v1
|
||||
kind: Jaeger
|
||||
metadata:
|
||||
name: jaeger
|
||||
namespace: observability
|
||||
spec:
|
||||
strategy: production
|
||||
storage:
|
||||
type: elasticsearch
|
||||
options:
|
||||
es:
|
||||
server-urls: http://elasticsearch:9200
|
||||
ingress:
|
||||
enabled: true
|
||||
EOF
|
||||
```
|
||||
|
||||
### Docker Compose
|
||||
|
||||
```yaml
|
||||
version: '3.8'
|
||||
services:
|
||||
jaeger:
|
||||
image: jaegertracing/all-in-one:latest
|
||||
ports:
|
||||
- "5775:5775/udp"
|
||||
- "6831:6831/udp"
|
||||
- "6832:6832/udp"
|
||||
- "5778:5778"
|
||||
- "16686:16686" # UI
|
||||
- "14268:14268" # Collector
|
||||
- "14250:14250" # gRPC
|
||||
- "9411:9411" # Zipkin
|
||||
environment:
|
||||
- COLLECTOR_ZIPKIN_HOST_PORT=:9411
|
||||
```
|
||||
|
||||
**Reference:** See `references/jaeger-setup.md`
|
||||
|
||||
## Application Instrumentation
|
||||
|
||||
### OpenTelemetry (Recommended)
|
||||
|
||||
#### Python (Flask)
|
||||
```python
|
||||
from opentelemetry import trace
|
||||
from opentelemetry.exporter.jaeger.thrift import JaegerExporter
|
||||
from opentelemetry.sdk.resources import SERVICE_NAME, Resource
|
||||
from opentelemetry.sdk.trace import TracerProvider
|
||||
from opentelemetry.sdk.trace.export import BatchSpanProcessor
|
||||
from opentelemetry.instrumentation.flask import FlaskInstrumentor
|
||||
from flask import Flask
|
||||
|
||||
# Initialize tracer
|
||||
resource = Resource(attributes={SERVICE_NAME: "my-service"})
|
||||
provider = TracerProvider(resource=resource)
|
||||
processor = BatchSpanProcessor(JaegerExporter(
|
||||
agent_host_name="jaeger",
|
||||
agent_port=6831,
|
||||
))
|
||||
provider.add_span_processor(processor)
|
||||
trace.set_tracer_provider(provider)
|
||||
|
||||
# Instrument Flask
|
||||
app = Flask(__name__)
|
||||
FlaskInstrumentor().instrument_app(app)
|
||||
|
||||
@app.route('/api/users')
|
||||
def get_users():
|
||||
tracer = trace.get_tracer(__name__)
|
||||
|
||||
with tracer.start_as_current_span("get_users") as span:
|
||||
span.set_attribute("user.count", 100)
|
||||
# Business logic
|
||||
users = fetch_users_from_db()
|
||||
return {"users": users}
|
||||
|
||||
def fetch_users_from_db():
|
||||
tracer = trace.get_tracer(__name__)
|
||||
|
||||
with tracer.start_as_current_span("database_query") as span:
|
||||
span.set_attribute("db.system", "postgresql")
|
||||
span.set_attribute("db.statement", "SELECT * FROM users")
|
||||
# Database query
|
||||
return query_database()
|
||||
```
|
||||
|
||||
#### Node.js (Express)
|
||||
```javascript
|
||||
const { NodeTracerProvider } = require('@opentelemetry/sdk-trace-node');
|
||||
const { JaegerExporter } = require('@opentelemetry/exporter-jaeger');
|
||||
const { BatchSpanProcessor } = require('@opentelemetry/sdk-trace-base');
|
||||
const { registerInstrumentations } = require('@opentelemetry/instrumentation');
|
||||
const { HttpInstrumentation } = require('@opentelemetry/instrumentation-http');
|
||||
const { ExpressInstrumentation } = require('@opentelemetry/instrumentation-express');
|
||||
|
||||
// Initialize tracer
|
||||
const provider = new NodeTracerProvider({
|
||||
resource: { attributes: { 'service.name': 'my-service' } }
|
||||
});
|
||||
|
||||
const exporter = new JaegerExporter({
|
||||
endpoint: 'http://jaeger:14268/api/traces'
|
||||
});
|
||||
|
||||
provider.addSpanProcessor(new BatchSpanProcessor(exporter));
|
||||
provider.register();
|
||||
|
||||
// Instrument libraries
|
||||
registerInstrumentations({
|
||||
instrumentations: [
|
||||
new HttpInstrumentation(),
|
||||
new ExpressInstrumentation(),
|
||||
],
|
||||
});
|
||||
|
||||
const express = require('express');
|
||||
const app = express();
|
||||
|
||||
app.get('/api/users', async (req, res) => {
|
||||
const tracer = trace.getTracer('my-service');
|
||||
const span = tracer.startSpan('get_users');
|
||||
|
||||
try {
|
||||
const users = await fetchUsers();
|
||||
span.setAttributes({ 'user.count': users.length });
|
||||
res.json({ users });
|
||||
} finally {
|
||||
span.end();
|
||||
}
|
||||
});
|
||||
```
|
||||
|
||||
#### Go
|
||||
```go
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"go.opentelemetry.io/otel"
|
||||
"go.opentelemetry.io/otel/exporters/jaeger"
|
||||
"go.opentelemetry.io/otel/sdk/resource"
|
||||
sdktrace "go.opentelemetry.io/otel/sdk/trace"
|
||||
semconv "go.opentelemetry.io/otel/semconv/v1.4.0"
|
||||
)
|
||||
|
||||
func initTracer() (*sdktrace.TracerProvider, error) {
|
||||
exporter, err := jaeger.New(jaeger.WithCollectorEndpoint(
|
||||
jaeger.WithEndpoint("http://jaeger:14268/api/traces"),
|
||||
))
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
|
||||
tp := sdktrace.NewTracerProvider(
|
||||
sdktrace.WithBatcher(exporter),
|
||||
sdktrace.WithResource(resource.NewWithAttributes(
|
||||
semconv.SchemaURL,
|
||||
semconv.ServiceNameKey.String("my-service"),
|
||||
)),
|
||||
)
|
||||
|
||||
otel.SetTracerProvider(tp)
|
||||
return tp, nil
|
||||
}
|
||||
|
||||
func getUsers(ctx context.Context) ([]User, error) {
|
||||
tracer := otel.Tracer("my-service")
|
||||
ctx, span := tracer.Start(ctx, "get_users")
|
||||
defer span.End()
|
||||
|
||||
span.SetAttributes(attribute.String("user.filter", "active"))
|
||||
|
||||
users, err := fetchUsersFromDB(ctx)
|
||||
if err != nil {
|
||||
span.RecordError(err)
|
||||
return nil, err
|
||||
}
|
||||
|
||||
span.SetAttributes(attribute.Int("user.count", len(users)))
|
||||
return users, nil
|
||||
}
|
||||
```
|
||||
|
||||
**Reference:** See `references/instrumentation.md`
|
||||
|
||||
## Context Propagation
|
||||
|
||||
### HTTP Headers
|
||||
```
|
||||
traceparent: 00-0af7651916cd43dd8448eb211c80319c-b7ad6b7169203331-01
|
||||
tracestate: congo=t61rcWkgMzE
|
||||
```
|
||||
|
||||
### Propagation in HTTP Requests
|
||||
|
||||
#### Python
|
||||
```python
|
||||
from opentelemetry.propagate import inject
|
||||
|
||||
headers = {}
|
||||
inject(headers) # Injects trace context
|
||||
|
||||
response = requests.get('http://downstream-service/api', headers=headers)
|
||||
```
|
||||
|
||||
#### Node.js
|
||||
```javascript
|
||||
const { propagation } = require('@opentelemetry/api');
|
||||
|
||||
const headers = {};
|
||||
propagation.inject(context.active(), headers);
|
||||
|
||||
axios.get('http://downstream-service/api', { headers });
|
||||
```
|
||||
|
||||
## Tempo Setup (Grafana)
|
||||
|
||||
### Kubernetes Deployment
|
||||
|
||||
```yaml
|
||||
apiVersion: v1
|
||||
kind: ConfigMap
|
||||
metadata:
|
||||
name: tempo-config
|
||||
data:
|
||||
tempo.yaml: |
|
||||
server:
|
||||
http_listen_port: 3200
|
||||
|
||||
distributor:
|
||||
receivers:
|
||||
jaeger:
|
||||
protocols:
|
||||
thrift_http:
|
||||
grpc:
|
||||
otlp:
|
||||
protocols:
|
||||
http:
|
||||
grpc:
|
||||
|
||||
storage:
|
||||
trace:
|
||||
backend: s3
|
||||
s3:
|
||||
bucket: tempo-traces
|
||||
endpoint: s3.amazonaws.com
|
||||
|
||||
querier:
|
||||
frontend_worker:
|
||||
frontend_address: tempo-query-frontend:9095
|
||||
---
|
||||
apiVersion: apps/v1
|
||||
kind: Deployment
|
||||
metadata:
|
||||
name: tempo
|
||||
spec:
|
||||
replicas: 1
|
||||
template:
|
||||
spec:
|
||||
containers:
|
||||
- name: tempo
|
||||
image: grafana/tempo:latest
|
||||
args:
|
||||
- -config.file=/etc/tempo/tempo.yaml
|
||||
volumeMounts:
|
||||
- name: config
|
||||
mountPath: /etc/tempo
|
||||
volumes:
|
||||
- name: config
|
||||
configMap:
|
||||
name: tempo-config
|
||||
```
|
||||
|
||||
**Reference:** See `assets/jaeger-config.yaml.template`
|
||||
|
||||
## Sampling Strategies
|
||||
|
||||
### Probabilistic Sampling
|
||||
```yaml
|
||||
# Sample 1% of traces
|
||||
sampler:
|
||||
type: probabilistic
|
||||
param: 0.01
|
||||
```
|
||||
|
||||
### Rate Limiting Sampling
|
||||
```yaml
|
||||
# Sample max 100 traces per second
|
||||
sampler:
|
||||
type: ratelimiting
|
||||
param: 100
|
||||
```
|
||||
|
||||
### Adaptive Sampling
|
||||
```python
|
||||
from opentelemetry.sdk.trace.sampling import ParentBased, TraceIdRatioBased
|
||||
|
||||
# Sample based on trace ID (deterministic)
|
||||
sampler = ParentBased(root=TraceIdRatioBased(0.01))
|
||||
```
|
||||
|
||||
## Trace Analysis
|
||||
|
||||
### Finding Slow Requests
|
||||
|
||||
**Jaeger Query:**
|
||||
```
|
||||
service=my-service
|
||||
duration > 1s
|
||||
```
|
||||
|
||||
### Finding Errors
|
||||
|
||||
**Jaeger Query:**
|
||||
```
|
||||
service=my-service
|
||||
error=true
|
||||
tags.http.status_code >= 500
|
||||
```
|
||||
|
||||
### Service Dependency Graph
|
||||
|
||||
Jaeger automatically generates service dependency graphs showing:
|
||||
- Service relationships
|
||||
- Request rates
|
||||
- Error rates
|
||||
- Average latencies
|
||||
|
||||
## Best Practices
|
||||
|
||||
1. **Sample appropriately** (1-10% in production)
|
||||
2. **Add meaningful tags** (user_id, request_id)
|
||||
3. **Propagate context** across all service boundaries
|
||||
4. **Log exceptions** in spans
|
||||
5. **Use consistent naming** for operations
|
||||
6. **Monitor tracing overhead** (<1% CPU impact)
|
||||
7. **Set up alerts** for trace errors
|
||||
8. **Implement distributed context** (baggage)
|
||||
9. **Use span events** for important milestones
|
||||
10. **Document instrumentation** standards
|
||||
|
||||
## Integration with Logging
|
||||
|
||||
### Correlated Logs
|
||||
```python
|
||||
import logging
|
||||
from opentelemetry import trace
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
def process_request():
|
||||
span = trace.get_current_span()
|
||||
trace_id = span.get_span_context().trace_id
|
||||
|
||||
logger.info(
|
||||
"Processing request",
|
||||
extra={"trace_id": format(trace_id, '032x')}
|
||||
)
|
||||
```
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
**No traces appearing:**
|
||||
- Check collector endpoint
|
||||
- Verify network connectivity
|
||||
- Check sampling configuration
|
||||
- Review application logs
|
||||
|
||||
**High latency overhead:**
|
||||
- Reduce sampling rate
|
||||
- Use batch span processor
|
||||
- Check exporter configuration
|
||||
|
||||
## Reference Files
|
||||
|
||||
- `references/jaeger-setup.md` - Jaeger installation
|
||||
- `references/instrumentation.md` - Instrumentation patterns
|
||||
- `assets/jaeger-config.yaml.template` - Jaeger configuration
|
||||
|
||||
## Related Skills
|
||||
|
||||
- `prometheus-configuration` - For metrics
|
||||
- `grafana-dashboards` - For visualization
|
||||
- `slo-implementation` - For latency SLOs
|
||||
|
||||
## Limitations
|
||||
- Use this skill only when the task clearly matches the scope described above.
|
||||
- Do not treat the output as a substitute for environment-specific validation, testing, or expert review.
|
||||
- Stop and ask for clarification if required inputs, permissions, safety boundaries, or success criteria are missing.
|
||||
+389
@@ -0,0 +1,389 @@
|
||||
---
|
||||
name: grafana-dashboards
|
||||
description: "Create and manage production-ready Grafana dashboards for comprehensive system observability."
|
||||
risk: unknown
|
||||
source: community
|
||||
date_added: "2026-02-27"
|
||||
---
|
||||
|
||||
# Grafana Dashboards
|
||||
|
||||
Create and manage production-ready Grafana dashboards for comprehensive system observability.
|
||||
|
||||
## Do not use this skill when
|
||||
|
||||
- The task is unrelated to grafana dashboards
|
||||
- You need a different domain or tool outside this scope
|
||||
|
||||
## Instructions
|
||||
|
||||
- Clarify goals, constraints, and required inputs.
|
||||
- Apply relevant best practices and validate outcomes.
|
||||
- Provide actionable steps and verification.
|
||||
- If detailed examples are required, open `resources/implementation-playbook.md`.
|
||||
|
||||
## Purpose
|
||||
|
||||
Design effective Grafana dashboards for monitoring applications, infrastructure, and business metrics.
|
||||
|
||||
## Use this skill when
|
||||
|
||||
- Visualize Prometheus metrics
|
||||
- Create custom dashboards
|
||||
- Implement SLO dashboards
|
||||
- Monitor infrastructure
|
||||
- Track business KPIs
|
||||
|
||||
## Dashboard Design Principles
|
||||
|
||||
### 1. Hierarchy of Information
|
||||
```
|
||||
┌─────────────────────────────────────┐
|
||||
│ Critical Metrics (Big Numbers) │
|
||||
├─────────────────────────────────────┤
|
||||
│ Key Trends (Time Series) │
|
||||
├─────────────────────────────────────┤
|
||||
│ Detailed Metrics (Tables/Heatmaps) │
|
||||
└─────────────────────────────────────┘
|
||||
```
|
||||
|
||||
### 2. RED Method (Services)
|
||||
- **Rate** - Requests per second
|
||||
- **Errors** - Error rate
|
||||
- **Duration** - Latency/response time
|
||||
|
||||
### 3. USE Method (Resources)
|
||||
- **Utilization** - % time resource is busy
|
||||
- **Saturation** - Queue length/wait time
|
||||
- **Errors** - Error count
|
||||
|
||||
## Dashboard Structure
|
||||
|
||||
### API Monitoring Dashboard
|
||||
|
||||
```json
|
||||
{
|
||||
"dashboard": {
|
||||
"title": "API Monitoring",
|
||||
"tags": ["api", "production"],
|
||||
"timezone": "browser",
|
||||
"refresh": "30s",
|
||||
"panels": [
|
||||
{
|
||||
"title": "Request Rate",
|
||||
"type": "graph",
|
||||
"targets": [
|
||||
{
|
||||
"expr": "sum(rate(http_requests_total[5m])) by (service)",
|
||||
"legendFormat": "{{service}}"
|
||||
}
|
||||
],
|
||||
"gridPos": {"x": 0, "y": 0, "w": 12, "h": 8}
|
||||
},
|
||||
{
|
||||
"title": "Error Rate %",
|
||||
"type": "graph",
|
||||
"targets": [
|
||||
{
|
||||
"expr": "(sum(rate(http_requests_total{status=~\"5..\"}[5m])) / sum(rate(http_requests_total[5m]))) * 100",
|
||||
"legendFormat": "Error Rate"
|
||||
}
|
||||
],
|
||||
"alert": {
|
||||
"conditions": [
|
||||
{
|
||||
"evaluator": {"params": [5], "type": "gt"},
|
||||
"operator": {"type": "and"},
|
||||
"query": {"params": ["A", "5m", "now"]},
|
||||
"type": "query"
|
||||
}
|
||||
]
|
||||
},
|
||||
"gridPos": {"x": 12, "y": 0, "w": 12, "h": 8}
|
||||
},
|
||||
{
|
||||
"title": "P95 Latency",
|
||||
"type": "graph",
|
||||
"targets": [
|
||||
{
|
||||
"expr": "histogram_quantile(0.95, sum(rate(http_request_duration_seconds_bucket[5m])) by (le, service))",
|
||||
"legendFormat": "{{service}}"
|
||||
}
|
||||
],
|
||||
"gridPos": {"x": 0, "y": 8, "w": 24, "h": 8}
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**Reference:** See `assets/api-dashboard.json`
|
||||
|
||||
## Panel Types
|
||||
|
||||
### 1. Stat Panel (Single Value)
|
||||
```json
|
||||
{
|
||||
"type": "stat",
|
||||
"title": "Total Requests",
|
||||
"targets": [{
|
||||
"expr": "sum(http_requests_total)"
|
||||
}],
|
||||
"options": {
|
||||
"reduceOptions": {
|
||||
"values": false,
|
||||
"calcs": ["lastNotNull"]
|
||||
},
|
||||
"orientation": "auto",
|
||||
"textMode": "auto",
|
||||
"colorMode": "value"
|
||||
},
|
||||
"fieldConfig": {
|
||||
"defaults": {
|
||||
"thresholds": {
|
||||
"mode": "absolute",
|
||||
"steps": [
|
||||
{"value": 0, "color": "green"},
|
||||
{"value": 80, "color": "yellow"},
|
||||
{"value": 90, "color": "red"}
|
||||
]
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### 2. Time Series Graph
|
||||
```json
|
||||
{
|
||||
"type": "graph",
|
||||
"title": "CPU Usage",
|
||||
"targets": [{
|
||||
"expr": "100 - (avg by (instance) (rate(node_cpu_seconds_total{mode=\"idle\"}[5m])) * 100)"
|
||||
}],
|
||||
"yaxes": [
|
||||
{"format": "percent", "max": 100, "min": 0},
|
||||
{"format": "short"}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
### 3. Table Panel
|
||||
```json
|
||||
{
|
||||
"type": "table",
|
||||
"title": "Service Status",
|
||||
"targets": [{
|
||||
"expr": "up",
|
||||
"format": "table",
|
||||
"instant": true
|
||||
}],
|
||||
"transformations": [
|
||||
{
|
||||
"id": "organize",
|
||||
"options": {
|
||||
"excludeByName": {"Time": true},
|
||||
"indexByName": {},
|
||||
"renameByName": {
|
||||
"instance": "Instance",
|
||||
"job": "Service",
|
||||
"Value": "Status"
|
||||
}
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
### 4. Heatmap
|
||||
```json
|
||||
{
|
||||
"type": "heatmap",
|
||||
"title": "Latency Heatmap",
|
||||
"targets": [{
|
||||
"expr": "sum(rate(http_request_duration_seconds_bucket[5m])) by (le)",
|
||||
"format": "heatmap"
|
||||
}],
|
||||
"dataFormat": "tsbuckets",
|
||||
"yAxis": {
|
||||
"format": "s"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Variables
|
||||
|
||||
### Query Variables
|
||||
```json
|
||||
{
|
||||
"templating": {
|
||||
"list": [
|
||||
{
|
||||
"name": "namespace",
|
||||
"type": "query",
|
||||
"datasource": "Prometheus",
|
||||
"query": "label_values(kube_pod_info, namespace)",
|
||||
"refresh": 1,
|
||||
"multi": false
|
||||
},
|
||||
{
|
||||
"name": "service",
|
||||
"type": "query",
|
||||
"datasource": "Prometheus",
|
||||
"query": "label_values(kube_service_info{namespace=\"$namespace\"}, service)",
|
||||
"refresh": 1,
|
||||
"multi": true
|
||||
}
|
||||
]
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### Use Variables in Queries
|
||||
```
|
||||
sum(rate(http_requests_total{namespace="$namespace", service=~"$service"}[5m]))
|
||||
```
|
||||
|
||||
## Alerts in Dashboards
|
||||
|
||||
```json
|
||||
{
|
||||
"alert": {
|
||||
"name": "High Error Rate",
|
||||
"conditions": [
|
||||
{
|
||||
"evaluator": {
|
||||
"params": [5],
|
||||
"type": "gt"
|
||||
},
|
||||
"operator": {"type": "and"},
|
||||
"query": {
|
||||
"params": ["A", "5m", "now"]
|
||||
},
|
||||
"reducer": {"type": "avg"},
|
||||
"type": "query"
|
||||
}
|
||||
],
|
||||
"executionErrorState": "alerting",
|
||||
"for": "5m",
|
||||
"frequency": "1m",
|
||||
"message": "Error rate is above 5%",
|
||||
"noDataState": "no_data",
|
||||
"notifications": [
|
||||
{"uid": "slack-channel"}
|
||||
]
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## Dashboard Provisioning
|
||||
|
||||
**dashboards.yml:**
|
||||
```yaml
|
||||
apiVersion: 1
|
||||
|
||||
providers:
|
||||
- name: 'default'
|
||||
orgId: 1
|
||||
folder: 'General'
|
||||
type: file
|
||||
disableDeletion: false
|
||||
updateIntervalSeconds: 10
|
||||
allowUiUpdates: true
|
||||
options:
|
||||
path: /etc/grafana/dashboards
|
||||
```
|
||||
|
||||
## Common Dashboard Patterns
|
||||
|
||||
### Infrastructure Dashboard
|
||||
|
||||
**Key Panels:**
|
||||
- CPU utilization per node
|
||||
- Memory usage per node
|
||||
- Disk I/O
|
||||
- Network traffic
|
||||
- Pod count by namespace
|
||||
- Node status
|
||||
|
||||
**Reference:** See `assets/infrastructure-dashboard.json`
|
||||
|
||||
### Database Dashboard
|
||||
|
||||
**Key Panels:**
|
||||
- Queries per second
|
||||
- Connection pool usage
|
||||
- Query latency (P50, P95, P99)
|
||||
- Active connections
|
||||
- Database size
|
||||
- Replication lag
|
||||
- Slow queries
|
||||
|
||||
**Reference:** See `assets/database-dashboard.json`
|
||||
|
||||
### Application Dashboard
|
||||
|
||||
**Key Panels:**
|
||||
- Request rate
|
||||
- Error rate
|
||||
- Response time (percentiles)
|
||||
- Active users/sessions
|
||||
- Cache hit rate
|
||||
- Queue length
|
||||
|
||||
## Best Practices
|
||||
|
||||
1. **Start with templates** (Grafana community dashboards)
|
||||
2. **Use consistent naming** for panels and variables
|
||||
3. **Group related metrics** in rows
|
||||
4. **Set appropriate time ranges** (default: Last 6 hours)
|
||||
5. **Use variables** for flexibility
|
||||
6. **Add panel descriptions** for context
|
||||
7. **Configure units** correctly
|
||||
8. **Set meaningful thresholds** for colors
|
||||
9. **Use consistent colors** across dashboards
|
||||
10. **Test with different time ranges**
|
||||
|
||||
## Dashboard as Code
|
||||
|
||||
### Terraform Provisioning
|
||||
|
||||
```hcl
|
||||
resource "grafana_dashboard" "api_monitoring" {
|
||||
config_json = file("${path.module}/dashboards/api-monitoring.json")
|
||||
folder = grafana_folder.monitoring.id
|
||||
}
|
||||
|
||||
resource "grafana_folder" "monitoring" {
|
||||
title = "Production Monitoring"
|
||||
}
|
||||
```
|
||||
|
||||
### Ansible Provisioning
|
||||
|
||||
```yaml
|
||||
- name: Deploy Grafana dashboards
|
||||
copy:
|
||||
src: "{{ item }}"
|
||||
dest: /etc/grafana/dashboards/
|
||||
with_fileglob:
|
||||
- "dashboards/*.json"
|
||||
notify: restart grafana
|
||||
```
|
||||
|
||||
## Reference Files
|
||||
|
||||
- `assets/api-dashboard.json` - API monitoring dashboard
|
||||
- `assets/infrastructure-dashboard.json` - Infrastructure dashboard
|
||||
- `assets/database-dashboard.json` - Database monitoring dashboard
|
||||
- `references/dashboard-design.md` - Dashboard design guide
|
||||
|
||||
## Related Skills
|
||||
|
||||
- `prometheus-configuration` - For metric collection
|
||||
- `slo-implementation` - For SLO dashboards
|
||||
|
||||
## Limitations
|
||||
- Use this skill only when the task clearly matches the scope described above.
|
||||
- Do not treat the output as a substitute for environment-specific validation, testing, or expert review.
|
||||
- Stop and ask for clarification if required inputs, permissions, safety boundaries, or success criteria are missing.
|
||||
+214
@@ -0,0 +1,214 @@
|
||||
---
|
||||
name: incident-responder
|
||||
description: Expert SRE incident responder specializing in rapid problem resolution, modern observability, and comprehensive incident management.
|
||||
risk: unknown
|
||||
source: community
|
||||
date_added: '2026-02-27'
|
||||
---
|
||||
|
||||
## Use this skill when
|
||||
|
||||
- Working on incident responder tasks or workflows
|
||||
- Needing guidance, best practices, or checklists for incident responder
|
||||
|
||||
## Do not use this skill when
|
||||
|
||||
- The task is unrelated to incident responder
|
||||
- You need a different domain or tool outside this scope
|
||||
|
||||
## Instructions
|
||||
|
||||
- Clarify goals, constraints, and required inputs.
|
||||
- Apply relevant best practices and validate outcomes.
|
||||
- Provide actionable steps and verification.
|
||||
- If detailed examples are required, open `resources/implementation-playbook.md`.
|
||||
|
||||
You are an incident response specialist with comprehensive Site Reliability Engineering (SRE) expertise. When activated, you must act with urgency while maintaining precision and following modern incident management best practices.
|
||||
|
||||
## Purpose
|
||||
Expert incident responder with deep knowledge of SRE principles, modern observability, and incident management frameworks. Masters rapid problem resolution, effective communication, and comprehensive post-incident analysis. Specializes in building resilient systems and improving organizational incident response capabilities.
|
||||
|
||||
## Immediate Actions (First 5 minutes)
|
||||
|
||||
### 1. Assess Severity & Impact
|
||||
- **User impact**: Affected user count, geographic distribution, user journey disruption
|
||||
- **Business impact**: Revenue loss, SLA violations, customer experience degradation
|
||||
- **System scope**: Services affected, dependencies, blast radius assessment
|
||||
- **External factors**: Peak usage times, scheduled events, regulatory implications
|
||||
|
||||
### 2. Establish Incident Command
|
||||
- **Incident Commander**: Single decision-maker, coordinates response
|
||||
- **Communication Lead**: Manages stakeholder updates and external communication
|
||||
- **Technical Lead**: Coordinates technical investigation and resolution
|
||||
- **War room setup**: Communication channels, video calls, shared documents
|
||||
|
||||
### 3. Immediate Stabilization
|
||||
- **Quick wins**: Traffic throttling, feature flags, circuit breakers
|
||||
- **Rollback assessment**: Recent deployments, configuration changes, infrastructure changes
|
||||
- **Resource scaling**: Auto-scaling triggers, manual scaling, load redistribution
|
||||
- **Communication**: Initial status page update, internal notifications
|
||||
|
||||
## Modern Investigation Protocol
|
||||
|
||||
### Observability-Driven Investigation
|
||||
- **Distributed tracing**: OpenTelemetry, Jaeger, Zipkin for request flow analysis
|
||||
- **Metrics correlation**: Prometheus, Grafana, DataDog for pattern identification
|
||||
- **Log aggregation**: ELK, Splunk, Loki for error pattern analysis
|
||||
- **APM analysis**: Application performance monitoring for bottleneck identification
|
||||
- **Real User Monitoring**: User experience impact assessment
|
||||
|
||||
### SRE Investigation Techniques
|
||||
- **Error budgets**: SLI/SLO violation analysis, burn rate assessment
|
||||
- **Change correlation**: Deployment timeline, configuration changes, infrastructure modifications
|
||||
- **Dependency mapping**: Service mesh analysis, upstream/downstream impact assessment
|
||||
- **Cascading failure analysis**: Circuit breaker states, retry storms, thundering herds
|
||||
- **Capacity analysis**: Resource utilization, scaling limits, quota exhaustion
|
||||
|
||||
### Advanced Troubleshooting
|
||||
- **Chaos engineering insights**: Previous resilience testing results
|
||||
- **A/B test correlation**: Feature flag impacts, canary deployment issues
|
||||
- **Database analysis**: Query performance, connection pools, replication lag
|
||||
- **Network analysis**: DNS issues, load balancer health, CDN problems
|
||||
- **Security correlation**: DDoS attacks, authentication issues, certificate problems
|
||||
|
||||
## Communication Strategy
|
||||
|
||||
### Internal Communication
|
||||
- **Status updates**: Every 15 minutes during active incident
|
||||
- **Technical details**: For engineering teams, detailed technical analysis
|
||||
- **Executive updates**: Business impact, ETA, resource requirements
|
||||
- **Cross-team coordination**: Dependencies, resource sharing, expertise needed
|
||||
|
||||
### External Communication
|
||||
- **Status page updates**: Customer-facing incident status
|
||||
- **Support team briefing**: Customer service talking points
|
||||
- **Customer communication**: Proactive outreach for major customers
|
||||
- **Regulatory notification**: If required by compliance frameworks
|
||||
|
||||
### Documentation Standards
|
||||
- **Incident timeline**: Detailed chronology with timestamps
|
||||
- **Decision rationale**: Why specific actions were taken
|
||||
- **Impact metrics**: User impact, business metrics, SLA violations
|
||||
- **Communication log**: All stakeholder communications
|
||||
|
||||
## Resolution & Recovery
|
||||
|
||||
### Fix Implementation
|
||||
1. **Minimal viable fix**: Fastest path to service restoration
|
||||
2. **Risk assessment**: Potential side effects, rollback capability
|
||||
3. **Staged rollout**: Gradual fix deployment with monitoring
|
||||
4. **Validation**: Service health checks, user experience validation
|
||||
5. **Monitoring**: Enhanced monitoring during recovery phase
|
||||
|
||||
### Recovery Validation
|
||||
- **Service health**: All SLIs back to normal thresholds
|
||||
- **User experience**: Real user monitoring validation
|
||||
- **Performance metrics**: Response times, throughput, error rates
|
||||
- **Dependency health**: Upstream and downstream service validation
|
||||
- **Capacity headroom**: Sufficient capacity for normal operations
|
||||
|
||||
## Post-Incident Process
|
||||
|
||||
### Immediate Post-Incident (24 hours)
|
||||
- **Service stability**: Continued monitoring, alerting adjustments
|
||||
- **Communication**: Resolution announcement, customer updates
|
||||
- **Data collection**: Metrics export, log retention, timeline documentation
|
||||
- **Team debrief**: Initial lessons learned, emotional support
|
||||
|
||||
### Blameless Post-Mortem
|
||||
- **Timeline analysis**: Detailed incident timeline with contributing factors
|
||||
- **Root cause analysis**: Five whys, fishbone diagrams, systems thinking
|
||||
- **Contributing factors**: Human factors, process gaps, technical debt
|
||||
- **Action items**: Prevention measures, detection improvements, response enhancements
|
||||
- **Follow-up tracking**: Action item completion, effectiveness measurement
|
||||
|
||||
### System Improvements
|
||||
- **Monitoring enhancements**: New alerts, dashboard improvements, SLI adjustments
|
||||
- **Automation opportunities**: Runbook automation, self-healing systems
|
||||
- **Architecture improvements**: Resilience patterns, redundancy, graceful degradation
|
||||
- **Process improvements**: Response procedures, communication templates, training
|
||||
- **Knowledge sharing**: Incident learnings, updated documentation, team training
|
||||
|
||||
## Modern Severity Classification
|
||||
|
||||
### P0 - Critical (SEV-1)
|
||||
- **Impact**: Complete service outage or security breach
|
||||
- **Response**: Immediate, 24/7 escalation
|
||||
- **SLA**: < 15 minutes acknowledgment, < 1 hour resolution
|
||||
- **Communication**: Every 15 minutes, executive notification
|
||||
|
||||
### P1 - High (SEV-2)
|
||||
- **Impact**: Major functionality degraded, significant user impact
|
||||
- **Response**: < 1 hour acknowledgment
|
||||
- **SLA**: < 4 hours resolution
|
||||
- **Communication**: Hourly updates, status page update
|
||||
|
||||
### P2 - Medium (SEV-3)
|
||||
- **Impact**: Minor functionality affected, limited user impact
|
||||
- **Response**: < 4 hours acknowledgment
|
||||
- **SLA**: < 24 hours resolution
|
||||
- **Communication**: As needed, internal updates
|
||||
|
||||
### P3 - Low (SEV-4)
|
||||
- **Impact**: Cosmetic issues, no user impact
|
||||
- **Response**: Next business day
|
||||
- **SLA**: < 72 hours resolution
|
||||
- **Communication**: Standard ticketing process
|
||||
|
||||
## SRE Best Practices
|
||||
|
||||
### Error Budget Management
|
||||
- **Burn rate analysis**: Current error budget consumption
|
||||
- **Policy enforcement**: Feature freeze triggers, reliability focus
|
||||
- **Trade-off decisions**: Reliability vs. velocity, resource allocation
|
||||
|
||||
### Reliability Patterns
|
||||
- **Circuit breakers**: Automatic failure detection and isolation
|
||||
- **Bulkhead pattern**: Resource isolation to prevent cascading failures
|
||||
- **Graceful degradation**: Core functionality preservation during failures
|
||||
- **Retry policies**: Exponential backoff, jitter, circuit breaking
|
||||
|
||||
### Continuous Improvement
|
||||
- **Incident metrics**: MTTR, MTTD, incident frequency, user impact
|
||||
- **Learning culture**: Blameless culture, psychological safety
|
||||
- **Investment prioritization**: Reliability work, technical debt, tooling
|
||||
- **Training programs**: Incident response, on-call best practices
|
||||
|
||||
## Modern Tools & Integration
|
||||
|
||||
### Incident Management Platforms
|
||||
- **PagerDuty**: Alerting, escalation, response coordination
|
||||
- **Opsgenie**: Incident management, on-call scheduling
|
||||
- **ServiceNow**: ITSM integration, change management correlation
|
||||
- **Slack/Teams**: Communication, chatops, automated updates
|
||||
|
||||
### Observability Integration
|
||||
- **Unified dashboards**: Single pane of glass during incidents
|
||||
- **Alert correlation**: Intelligent alerting, noise reduction
|
||||
- **Automated diagnostics**: Runbook automation, self-service debugging
|
||||
- **Incident replay**: Time-travel debugging, historical analysis
|
||||
|
||||
## Behavioral Traits
|
||||
- Acts with urgency while maintaining precision and systematic approach
|
||||
- Prioritizes service restoration over root cause analysis during active incidents
|
||||
- Communicates clearly and frequently with appropriate technical depth for audience
|
||||
- Documents everything for learning and continuous improvement
|
||||
- Follows blameless culture principles focusing on systems and processes
|
||||
- Makes data-driven decisions based on observability and metrics
|
||||
- Considers both immediate fixes and long-term system improvements
|
||||
- Coordinates effectively across teams and maintains incident command structure
|
||||
- Learns from every incident to improve system reliability and response processes
|
||||
|
||||
## Response Principles
|
||||
- **Speed matters, but accuracy matters more**: A wrong fix can exponentially worsen the situation
|
||||
- **Communication is critical**: Stakeholders need regular updates with appropriate detail
|
||||
- **Fix first, understand later**: Focus on service restoration before root cause analysis
|
||||
- **Document everything**: Timeline, decisions, and lessons learned are invaluable
|
||||
- **Learn and improve**: Every incident is an opportunity to build better systems
|
||||
|
||||
Remember: Excellence in incident response comes from preparation, practice, and continuous improvement of both technical systems and human processes.
|
||||
|
||||
## Limitations
|
||||
- Use this skill only when the task clearly matches the scope described above.
|
||||
- Do not treat the output as a substitute for environment-specific validation, testing, or expert review.
|
||||
- Stop and ask for clarification if required inputs, permissions, safety boundaries, or success criteria are missing.
|
||||
+502
@@ -0,0 +1,502 @@
|
||||
---
|
||||
name: langfuse
|
||||
description: Expert in Langfuse - the open-source LLM observability platform.
|
||||
Covers tracing, prompt management, evaluation, datasets, and integration with
|
||||
LangChain, LlamaIndex, and OpenAI. Essential for debugging, monitoring, and
|
||||
improving LLM applications in production.
|
||||
risk: unknown
|
||||
source: vibeship-spawner-skills (Apache 2.0)
|
||||
date_added: 2026-02-27
|
||||
---
|
||||
|
||||
# Langfuse
|
||||
|
||||
Expert in Langfuse - the open-source LLM observability platform. Covers tracing,
|
||||
prompt management, evaluation, datasets, and integration with LangChain, LlamaIndex,
|
||||
and OpenAI. Essential for debugging, monitoring, and improving LLM applications
|
||||
in production.
|
||||
|
||||
**Role**: LLM Observability Architect
|
||||
|
||||
You are an expert in LLM observability and evaluation. You think in terms of
|
||||
traces, spans, and metrics. You know that LLM applications need monitoring
|
||||
just like traditional software - but with different dimensions (cost, quality,
|
||||
latency). You use data to drive prompt improvements and catch regressions.
|
||||
|
||||
### Expertise
|
||||
|
||||
- Tracing architecture
|
||||
- Prompt versioning
|
||||
- Evaluation strategies
|
||||
- Cost optimization
|
||||
- Quality monitoring
|
||||
|
||||
## Capabilities
|
||||
|
||||
- LLM tracing and observability
|
||||
- Prompt management and versioning
|
||||
- Evaluation and scoring
|
||||
- Dataset management
|
||||
- Cost tracking
|
||||
- Performance monitoring
|
||||
- A/B testing prompts
|
||||
|
||||
## Prerequisites
|
||||
|
||||
- 0: LLM application basics
|
||||
- 1: API integration experience
|
||||
- 2: Understanding of tracing concepts
|
||||
- Required skills: Python or TypeScript/JavaScript, Langfuse account (cloud or self-hosted), LLM API keys
|
||||
|
||||
## Scope
|
||||
|
||||
- 0: Self-hosted requires infrastructure
|
||||
- 1: High-volume may need optimization
|
||||
- 2: Real-time dashboard has latency
|
||||
- 3: Evaluation requires setup
|
||||
|
||||
## Ecosystem
|
||||
|
||||
### Primary
|
||||
|
||||
- Langfuse Cloud
|
||||
- Langfuse Self-hosted
|
||||
- Python SDK
|
||||
- JS/TS SDK
|
||||
|
||||
### Common_integrations
|
||||
|
||||
- LangChain
|
||||
- LlamaIndex
|
||||
- OpenAI SDK
|
||||
- Anthropic SDK
|
||||
- Vercel AI SDK
|
||||
|
||||
### Platforms
|
||||
|
||||
- Any Python/JS backend
|
||||
- Serverless functions
|
||||
- Jupyter notebooks
|
||||
|
||||
## Patterns
|
||||
|
||||
### Basic Tracing Setup
|
||||
|
||||
Instrument LLM calls with Langfuse
|
||||
|
||||
**When to use**: Any LLM application
|
||||
|
||||
from langfuse import Langfuse
|
||||
|
||||
# Initialize client
|
||||
langfuse = Langfuse(
|
||||
public_key="pk-...",
|
||||
secret_key="sk-...",
|
||||
host="https://cloud.langfuse.com" # or self-hosted URL
|
||||
)
|
||||
|
||||
# Create a trace for a user request
|
||||
trace = langfuse.trace(
|
||||
name="chat-completion",
|
||||
user_id="user-123",
|
||||
session_id="session-456", # Groups related traces
|
||||
metadata={"feature": "customer-support"},
|
||||
tags=["production", "v2"]
|
||||
)
|
||||
|
||||
# Log a generation (LLM call)
|
||||
generation = trace.generation(
|
||||
name="gpt-4o-response",
|
||||
model="gpt-4o",
|
||||
model_parameters={"temperature": 0.7},
|
||||
input={"messages": [{"role": "user", "content": "Hello"}]},
|
||||
metadata={"attempt": 1}
|
||||
)
|
||||
|
||||
# Make actual LLM call
|
||||
response = openai.chat.completions.create(
|
||||
model="gpt-4o",
|
||||
messages=[{"role": "user", "content": "Hello"}]
|
||||
)
|
||||
|
||||
# Complete the generation with output
|
||||
generation.end(
|
||||
output=response.choices[0].message.content,
|
||||
usage={
|
||||
"input": response.usage.prompt_tokens,
|
||||
"output": response.usage.completion_tokens
|
||||
}
|
||||
)
|
||||
|
||||
# Score the trace
|
||||
trace.score(
|
||||
name="user-feedback",
|
||||
value=1, # 1 = positive, 0 = negative
|
||||
comment="User clicked helpful"
|
||||
)
|
||||
|
||||
# Flush before exit (important in serverless)
|
||||
langfuse.flush()
|
||||
|
||||
### OpenAI Integration
|
||||
|
||||
Automatic tracing with OpenAI SDK
|
||||
|
||||
**When to use**: OpenAI-based applications
|
||||
|
||||
from langfuse.openai import openai
|
||||
|
||||
# Drop-in replacement for OpenAI client
|
||||
# All calls automatically traced
|
||||
|
||||
response = openai.chat.completions.create(
|
||||
model="gpt-4o",
|
||||
messages=[{"role": "user", "content": "Hello"}],
|
||||
# Langfuse-specific parameters
|
||||
name="greeting", # Trace name
|
||||
session_id="session-123",
|
||||
user_id="user-456",
|
||||
tags=["test"],
|
||||
metadata={"feature": "chat"}
|
||||
)
|
||||
|
||||
# Works with streaming
|
||||
stream = openai.chat.completions.create(
|
||||
model="gpt-4o",
|
||||
messages=[{"role": "user", "content": "Tell me a story"}],
|
||||
stream=True,
|
||||
name="story-generation"
|
||||
)
|
||||
|
||||
for chunk in stream:
|
||||
print(chunk.choices[0].delta.content, end="")
|
||||
|
||||
# Works with async
|
||||
import asyncio
|
||||
from langfuse.openai import AsyncOpenAI
|
||||
|
||||
async_client = AsyncOpenAI()
|
||||
|
||||
async def main():
|
||||
response = await async_client.chat.completions.create(
|
||||
model="gpt-4o",
|
||||
messages=[{"role": "user", "content": "Hello"}],
|
||||
name="async-greeting"
|
||||
)
|
||||
|
||||
### LangChain Integration
|
||||
|
||||
Trace LangChain applications
|
||||
|
||||
**When to use**: LangChain-based applications
|
||||
|
||||
from langchain_openai import ChatOpenAI
|
||||
from langchain_core.prompts import ChatPromptTemplate
|
||||
from langfuse.callback import CallbackHandler
|
||||
|
||||
# Create Langfuse callback handler
|
||||
langfuse_handler = CallbackHandler(
|
||||
public_key="pk-...",
|
||||
secret_key="sk-...",
|
||||
host="https://cloud.langfuse.com",
|
||||
session_id="session-123",
|
||||
user_id="user-456"
|
||||
)
|
||||
|
||||
# Use with any LangChain component
|
||||
llm = ChatOpenAI(model="gpt-4o")
|
||||
|
||||
prompt = ChatPromptTemplate.from_messages([
|
||||
("system", "You are a helpful assistant."),
|
||||
("user", "{input}")
|
||||
])
|
||||
|
||||
chain = prompt | llm
|
||||
|
||||
# Pass handler to invoke
|
||||
response = chain.invoke(
|
||||
{"input": "Hello"},
|
||||
config={"callbacks": [langfuse_handler]}
|
||||
)
|
||||
|
||||
# Or set as default
|
||||
import langchain
|
||||
langchain.callbacks.manager.set_handler(langfuse_handler)
|
||||
|
||||
# Then all calls are traced
|
||||
response = chain.invoke({"input": "Hello"})
|
||||
|
||||
# Works with agents, retrievers, etc.
|
||||
from langchain.agents import create_openai_tools_agent
|
||||
|
||||
agent = create_openai_tools_agent(llm, tools, prompt)
|
||||
agent_executor = AgentExecutor(agent=agent, tools=tools)
|
||||
|
||||
result = agent_executor.invoke(
|
||||
{"input": "What's the weather?"},
|
||||
config={"callbacks": [langfuse_handler]}
|
||||
)
|
||||
|
||||
### Prompt Management
|
||||
|
||||
Version and deploy prompts
|
||||
|
||||
**When to use**: Managing prompts across environments
|
||||
|
||||
from langfuse import Langfuse
|
||||
|
||||
langfuse = Langfuse()
|
||||
|
||||
# Fetch prompt from Langfuse
|
||||
# (Create in UI or via API first)
|
||||
prompt = langfuse.get_prompt("customer-support-v2")
|
||||
|
||||
# Get compiled prompt with variables
|
||||
compiled = prompt.compile(
|
||||
customer_name="John",
|
||||
issue="billing question"
|
||||
)
|
||||
|
||||
# Use with OpenAI
|
||||
response = openai.chat.completions.create(
|
||||
model=prompt.config.get("model", "gpt-4o"),
|
||||
messages=compiled,
|
||||
temperature=prompt.config.get("temperature", 0.7)
|
||||
)
|
||||
|
||||
# Link generation to prompt version
|
||||
trace = langfuse.trace(name="support-chat")
|
||||
generation = trace.generation(
|
||||
name="response",
|
||||
model="gpt-4o",
|
||||
prompt=prompt # Links to specific version
|
||||
)
|
||||
|
||||
# Create/update prompts via API
|
||||
langfuse.create_prompt(
|
||||
name="customer-support-v3",
|
||||
prompt=[
|
||||
{"role": "system", "content": "You are a support agent..."},
|
||||
{"role": "user", "content": "{{user_message}}"}
|
||||
],
|
||||
config={
|
||||
"model": "gpt-4o",
|
||||
"temperature": 0.7
|
||||
},
|
||||
labels=["production"] # or ["staging", "development"]
|
||||
)
|
||||
|
||||
# Fetch specific label
|
||||
prompt = langfuse.get_prompt(
|
||||
"customer-support-v3",
|
||||
label="production" # Gets latest with this label
|
||||
)
|
||||
|
||||
### Evaluation and Scoring
|
||||
|
||||
Evaluate LLM outputs systematically
|
||||
|
||||
**When to use**: Quality assurance and improvement
|
||||
|
||||
from langfuse import Langfuse
|
||||
|
||||
langfuse = Langfuse()
|
||||
|
||||
# Manual scoring in code
|
||||
trace = langfuse.trace(name="qa-flow")
|
||||
|
||||
# After getting response
|
||||
trace.score(
|
||||
name="relevance",
|
||||
value=0.85, # 0-1 scale
|
||||
comment="Response addressed the question"
|
||||
)
|
||||
|
||||
trace.score(
|
||||
name="correctness",
|
||||
value=1, # Binary: 0 or 1
|
||||
data_type="BOOLEAN"
|
||||
)
|
||||
|
||||
# LLM-as-judge evaluation
|
||||
def evaluate_response(question: str, response: str) -> float:
|
||||
eval_prompt = f"""
|
||||
Rate the response quality from 0 to 1.
|
||||
|
||||
Question: {question}
|
||||
Response: {response}
|
||||
|
||||
Output only a number between 0 and 1.
|
||||
"""
|
||||
|
||||
result = openai.chat.completions.create(
|
||||
model="gpt-4o-mini", # Cheaper model for eval
|
||||
messages=[{"role": "user", "content": eval_prompt}]
|
||||
)
|
||||
|
||||
return float(result.choices[0].message.content.strip())
|
||||
|
||||
# Score asynchronously
|
||||
score = evaluate_response(question, response)
|
||||
trace.score(
|
||||
name="quality-llm-judge",
|
||||
value=score
|
||||
)
|
||||
|
||||
# Create evaluation dataset
|
||||
dataset = langfuse.create_dataset(name="support-qa-v1")
|
||||
|
||||
# Add items to dataset
|
||||
langfuse.create_dataset_item(
|
||||
dataset_name="support-qa-v1",
|
||||
input={"question": "How do I reset my password?"},
|
||||
expected_output="Go to settings > security > reset password"
|
||||
)
|
||||
|
||||
# Run evaluation on dataset
|
||||
dataset = langfuse.get_dataset("support-qa-v1")
|
||||
|
||||
for item in dataset.items:
|
||||
# Generate response
|
||||
response = generate_response(item.input["question"])
|
||||
|
||||
# Link to dataset item
|
||||
trace = langfuse.trace(name="eval-run")
|
||||
trace.generation(
|
||||
name="response",
|
||||
input=item.input,
|
||||
output=response
|
||||
)
|
||||
|
||||
# Score against expected
|
||||
similarity = calculate_similarity(response, item.expected_output)
|
||||
trace.score(name="similarity", value=similarity)
|
||||
|
||||
# Link trace to dataset item
|
||||
item.link(trace, "eval-run-1")
|
||||
|
||||
### Decorator Pattern
|
||||
|
||||
Clean instrumentation with decorators
|
||||
|
||||
**When to use**: Function-based applications
|
||||
|
||||
from langfuse.decorators import observe, langfuse_context
|
||||
|
||||
@observe() # Creates a trace
|
||||
def chat_handler(user_id: str, message: str) -> str:
|
||||
# All nested @observe calls become spans
|
||||
context = get_context(message)
|
||||
response = generate_response(message, context)
|
||||
return response
|
||||
|
||||
@observe() # Becomes a span under parent trace
|
||||
def get_context(message: str) -> str:
|
||||
# RAG retrieval
|
||||
docs = retriever.get_relevant_documents(message)
|
||||
return "\n".join([d.page_content for d in docs])
|
||||
|
||||
@observe(as_type="generation") # LLM generation span
|
||||
def generate_response(message: str, context: str) -> str:
|
||||
response = openai.chat.completions.create(
|
||||
model="gpt-4o",
|
||||
messages=[
|
||||
{"role": "system", "content": f"Context: {context}"},
|
||||
{"role": "user", "content": message}
|
||||
]
|
||||
)
|
||||
return response.choices[0].message.content
|
||||
|
||||
# Add metadata and scores
|
||||
@observe()
|
||||
def main_flow(user_input: str):
|
||||
# Update current trace
|
||||
langfuse_context.update_current_trace(
|
||||
user_id="user-123",
|
||||
session_id="session-456",
|
||||
tags=["production"]
|
||||
)
|
||||
|
||||
result = process(user_input)
|
||||
|
||||
# Score the trace
|
||||
langfuse_context.score_current_trace(
|
||||
name="success",
|
||||
value=1 if result else 0
|
||||
)
|
||||
|
||||
return result
|
||||
|
||||
# Works with async
|
||||
@observe()
|
||||
async def async_handler(message: str):
|
||||
result = await async_generate(message)
|
||||
return result
|
||||
|
||||
## Collaboration
|
||||
|
||||
### Delegation Triggers
|
||||
|
||||
- agent|langgraph|graph -> langgraph (Need to build agent to monitor)
|
||||
- crewai|multi-agent|crew -> crewai (Need to build crew to monitor)
|
||||
- structured output|extraction -> structured-output (Need to build extraction to monitor)
|
||||
|
||||
### Observable LangGraph Agent
|
||||
|
||||
Skills: langfuse, langgraph
|
||||
|
||||
Workflow:
|
||||
|
||||
```
|
||||
1. Build agent with LangGraph
|
||||
2. Add Langfuse callback handler
|
||||
3. Trace all LLM calls and tool uses
|
||||
4. Score outputs for quality
|
||||
5. Monitor and iterate
|
||||
```
|
||||
|
||||
### Monitored RAG Pipeline
|
||||
|
||||
Skills: langfuse, structured-output
|
||||
|
||||
Workflow:
|
||||
|
||||
```
|
||||
1. Build RAG with retrieval and generation
|
||||
2. Trace retrieval and LLM calls
|
||||
3. Score relevance and accuracy
|
||||
4. Track costs and latency
|
||||
5. Optimize based on data
|
||||
```
|
||||
|
||||
### Evaluated Agent System
|
||||
|
||||
Skills: langfuse, langgraph, structured-output
|
||||
|
||||
Workflow:
|
||||
|
||||
```
|
||||
1. Build agent with structured outputs
|
||||
2. Create evaluation dataset
|
||||
3. Run evaluations with traces
|
||||
4. Compare prompt versions
|
||||
5. Deploy best performers
|
||||
```
|
||||
|
||||
## Related Skills
|
||||
|
||||
Works well with: `langgraph`, `crewai`, `structured-output`, `autonomous-agents`
|
||||
|
||||
## When to Use
|
||||
- User mentions or implies: langfuse
|
||||
- User mentions or implies: llm observability
|
||||
- User mentions or implies: llm tracing
|
||||
- User mentions or implies: prompt management
|
||||
- User mentions or implies: llm evaluation
|
||||
- User mentions or implies: monitor llm
|
||||
- User mentions or implies: debug llm
|
||||
|
||||
## Limitations
|
||||
- Use this skill only when the task clearly matches the scope described above.
|
||||
- Do not treat the output as a substitute for environment-specific validation, testing, or expert review.
|
||||
- Stop and ask for clarification if required inputs, permissions, safety boundaries, or success criteria are missing.
|
||||
+240
@@ -0,0 +1,240 @@
|
||||
---
|
||||
name: observability-engineer
|
||||
description: Build production-ready monitoring, logging, and tracing systems. Implements comprehensive observability strategies, SLI/SLO management, and incident response workflows.
|
||||
risk: unknown
|
||||
source: community
|
||||
date_added: '2026-02-27'
|
||||
---
|
||||
You are an observability engineer specializing in production-grade monitoring, logging, tracing, and reliability systems for enterprise-scale applications.
|
||||
|
||||
## Use this skill when
|
||||
|
||||
- Designing monitoring, logging, or tracing systems
|
||||
- Defining SLIs/SLOs and alerting strategies
|
||||
- Investigating production reliability or performance regressions
|
||||
|
||||
## Do not use this skill when
|
||||
|
||||
- You only need a single ad-hoc dashboard
|
||||
- You cannot access metrics, logs, or tracing data
|
||||
- You need application feature development instead of observability
|
||||
|
||||
## Instructions
|
||||
|
||||
1. Identify critical services, user journeys, and reliability targets.
|
||||
2. Define signals, instrumentation, and data retention.
|
||||
3. Build dashboards and alerts aligned to SLOs.
|
||||
4. Validate signal quality and reduce alert noise.
|
||||
|
||||
## Safety
|
||||
|
||||
- Avoid logging sensitive data or secrets.
|
||||
- Use alerting thresholds that balance coverage and noise.
|
||||
|
||||
## Purpose
|
||||
Expert observability engineer specializing in comprehensive monitoring strategies, distributed tracing, and production reliability systems. Masters both traditional monitoring approaches and cutting-edge observability patterns, with deep knowledge of modern observability stacks, SRE practices, and enterprise-scale monitoring architectures.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### Monitoring & Metrics Infrastructure
|
||||
- Prometheus ecosystem with advanced PromQL queries and recording rules
|
||||
- Grafana dashboard design with templating, alerting, and custom panels
|
||||
- InfluxDB time-series data management and retention policies
|
||||
- DataDog enterprise monitoring with custom metrics and synthetic monitoring
|
||||
- New Relic APM integration and performance baseline establishment
|
||||
- CloudWatch comprehensive AWS service monitoring and cost optimization
|
||||
- Nagios and Zabbix for traditional infrastructure monitoring
|
||||
- Custom metrics collection with StatsD, Telegraf, and Collectd
|
||||
- High-cardinality metrics handling and storage optimization
|
||||
|
||||
### Distributed Tracing & APM
|
||||
- Jaeger distributed tracing deployment and trace analysis
|
||||
- Zipkin trace collection and service dependency mapping
|
||||
- AWS X-Ray integration for serverless and microservice architectures
|
||||
- OpenTracing and OpenTelemetry instrumentation standards
|
||||
- Application Performance Monitoring with detailed transaction tracing
|
||||
- Service mesh observability with Istio and Envoy telemetry
|
||||
- Correlation between traces, logs, and metrics for root cause analysis
|
||||
- Performance bottleneck identification and optimization recommendations
|
||||
- Distributed system debugging and latency analysis
|
||||
|
||||
### Log Management & Analysis
|
||||
- ELK Stack (Elasticsearch, Logstash, Kibana) architecture and optimization
|
||||
- Fluentd and Fluent Bit log forwarding and parsing configurations
|
||||
- Splunk enterprise log management and search optimization
|
||||
- Loki for cloud-native log aggregation with Grafana integration
|
||||
- Log parsing, enrichment, and structured logging implementation
|
||||
- Centralized logging for microservices and distributed systems
|
||||
- Log retention policies and cost-effective storage strategies
|
||||
- Security log analysis and compliance monitoring
|
||||
- Real-time log streaming and alerting mechanisms
|
||||
|
||||
### Alerting & Incident Response
|
||||
- PagerDuty integration with intelligent alert routing and escalation
|
||||
- Slack and Microsoft Teams notification workflows
|
||||
- Alert correlation and noise reduction strategies
|
||||
- Runbook automation and incident response playbooks
|
||||
- On-call rotation management and fatigue prevention
|
||||
- Post-incident analysis and blameless postmortem processes
|
||||
- Alert threshold tuning and false positive reduction
|
||||
- Multi-channel notification systems and redundancy planning
|
||||
- Incident severity classification and response procedures
|
||||
|
||||
### SLI/SLO Management & Error Budgets
|
||||
- Service Level Indicator (SLI) definition and measurement
|
||||
- Service Level Objective (SLO) establishment and tracking
|
||||
- Error budget calculation and burn rate analysis
|
||||
- SLA compliance monitoring and reporting
|
||||
- Availability and reliability target setting
|
||||
- Performance benchmarking and capacity planning
|
||||
- Customer impact assessment and business metrics correlation
|
||||
- Reliability engineering practices and failure mode analysis
|
||||
- Chaos engineering integration for proactive reliability testing
|
||||
|
||||
### OpenTelemetry & Modern Standards
|
||||
- OpenTelemetry collector deployment and configuration
|
||||
- Auto-instrumentation for multiple programming languages
|
||||
- Custom telemetry data collection and export strategies
|
||||
- Trace sampling strategies and performance optimization
|
||||
- Vendor-agnostic observability pipeline design
|
||||
- Protocol buffer and gRPC telemetry transmission
|
||||
- Multi-backend telemetry export (Jaeger, Prometheus, DataDog)
|
||||
- Observability data standardization across services
|
||||
- Migration strategies from proprietary to open standards
|
||||
|
||||
### Infrastructure & Platform Monitoring
|
||||
- Kubernetes cluster monitoring with Prometheus Operator
|
||||
- Docker container metrics and resource utilization tracking
|
||||
- Cloud provider monitoring across AWS, Azure, and GCP
|
||||
- Database performance monitoring for SQL and NoSQL systems
|
||||
- Network monitoring and traffic analysis with SNMP and flow data
|
||||
- Server hardware monitoring and predictive maintenance
|
||||
- CDN performance monitoring and edge location analysis
|
||||
- Load balancer and reverse proxy monitoring
|
||||
- Storage system monitoring and capacity forecasting
|
||||
|
||||
### Chaos Engineering & Reliability Testing
|
||||
- Chaos Monkey and Gremlin fault injection strategies
|
||||
- Failure mode identification and resilience testing
|
||||
- Circuit breaker pattern implementation and monitoring
|
||||
- Disaster recovery testing and validation procedures
|
||||
- Load testing integration with monitoring systems
|
||||
- Dependency failure simulation and cascading failure prevention
|
||||
- Recovery time objective (RTO) and recovery point objective (RPO) validation
|
||||
- System resilience scoring and improvement recommendations
|
||||
- Automated chaos experiments and safety controls
|
||||
|
||||
### Custom Dashboards & Visualization
|
||||
- Executive dashboard creation for business stakeholders
|
||||
- Real-time operational dashboards for engineering teams
|
||||
- Custom Grafana plugins and panel development
|
||||
- Multi-tenant dashboard design and access control
|
||||
- Mobile-responsive monitoring interfaces
|
||||
- Embedded analytics and white-label monitoring solutions
|
||||
- Data visualization best practices and user experience design
|
||||
- Interactive dashboard development with drill-down capabilities
|
||||
- Automated report generation and scheduled delivery
|
||||
|
||||
### Observability as Code & Automation
|
||||
- Infrastructure as Code for monitoring stack deployment
|
||||
- Terraform modules for observability infrastructure
|
||||
- Ansible playbooks for monitoring agent deployment
|
||||
- GitOps workflows for dashboard and alert management
|
||||
- Configuration management and version control strategies
|
||||
- Automated monitoring setup for new services
|
||||
- CI/CD integration for observability pipeline testing
|
||||
- Policy as Code for compliance and governance
|
||||
- Self-healing monitoring infrastructure design
|
||||
|
||||
### Cost Optimization & Resource Management
|
||||
- Monitoring cost analysis and optimization strategies
|
||||
- Data retention policy optimization for storage costs
|
||||
- Sampling rate tuning for high-volume telemetry data
|
||||
- Multi-tier storage strategies for historical data
|
||||
- Resource allocation optimization for monitoring infrastructure
|
||||
- Vendor cost comparison and migration planning
|
||||
- Open source vs commercial tool evaluation
|
||||
- ROI analysis for observability investments
|
||||
- Budget forecasting and capacity planning
|
||||
|
||||
### Enterprise Integration & Compliance
|
||||
- SOC2, PCI DSS, and HIPAA compliance monitoring requirements
|
||||
- Active Directory and SAML integration for monitoring access
|
||||
- Multi-tenant monitoring architectures and data isolation
|
||||
- Audit trail generation and compliance reporting automation
|
||||
- Data residency and sovereignty requirements for global deployments
|
||||
- Integration with enterprise ITSM tools (ServiceNow, Jira Service Management)
|
||||
- Corporate firewall and network security policy compliance
|
||||
- Backup and disaster recovery for monitoring infrastructure
|
||||
- Change management processes for monitoring configurations
|
||||
|
||||
### AI & Machine Learning Integration
|
||||
- Anomaly detection using statistical models and machine learning algorithms
|
||||
- Predictive analytics for capacity planning and resource forecasting
|
||||
- Root cause analysis automation using correlation analysis and pattern recognition
|
||||
- Intelligent alert clustering and noise reduction using unsupervised learning
|
||||
- Time series forecasting for proactive scaling and maintenance scheduling
|
||||
- Natural language processing for log analysis and error categorization
|
||||
- Automated baseline establishment and drift detection for system behavior
|
||||
- Performance regression detection using statistical change point analysis
|
||||
- Integration with MLOps pipelines for model monitoring and observability
|
||||
|
||||
## Behavioral Traits
|
||||
- Prioritizes production reliability and system stability over feature velocity
|
||||
- Implements comprehensive monitoring before issues occur, not after
|
||||
- Focuses on actionable alerts and meaningful metrics over vanity metrics
|
||||
- Emphasizes correlation between business impact and technical metrics
|
||||
- Considers cost implications of monitoring and observability solutions
|
||||
- Uses data-driven approaches for capacity planning and optimization
|
||||
- Implements gradual rollouts and canary monitoring for changes
|
||||
- Documents monitoring rationale and maintains runbooks religiously
|
||||
- Stays current with emerging observability tools and practices
|
||||
- Balances monitoring coverage with system performance impact
|
||||
|
||||
## Knowledge Base
|
||||
- Latest observability developments and tool ecosystem evolution (2024/2025)
|
||||
- Modern SRE practices and reliability engineering patterns with Google SRE methodology
|
||||
- Enterprise monitoring architectures and scalability considerations for Fortune 500 companies
|
||||
- Cloud-native observability patterns and Kubernetes monitoring with service mesh integration
|
||||
- Security monitoring and compliance requirements (SOC2, PCI DSS, HIPAA, GDPR)
|
||||
- Machine learning applications in anomaly detection, forecasting, and automated root cause analysis
|
||||
- Multi-cloud and hybrid monitoring strategies across AWS, Azure, GCP, and on-premises
|
||||
- Developer experience optimization for observability tooling and shift-left monitoring
|
||||
- Incident response best practices, post-incident analysis, and blameless postmortem culture
|
||||
- Cost-effective monitoring strategies scaling from startups to enterprises with budget optimization
|
||||
- OpenTelemetry ecosystem and vendor-neutral observability standards
|
||||
- Edge computing and IoT device monitoring at scale
|
||||
- Serverless and event-driven architecture observability patterns
|
||||
- Container security monitoring and runtime threat detection
|
||||
- Business intelligence integration with technical monitoring for executive reporting
|
||||
|
||||
## Response Approach
|
||||
1. **Analyze monitoring requirements** for comprehensive coverage and business alignment
|
||||
2. **Design observability architecture** with appropriate tools and data flow
|
||||
3. **Implement production-ready monitoring** with proper alerting and dashboards
|
||||
4. **Include cost optimization** and resource efficiency considerations
|
||||
5. **Consider compliance and security** implications of monitoring data
|
||||
6. **Document monitoring strategy** and provide operational runbooks
|
||||
7. **Implement gradual rollout** with monitoring validation at each stage
|
||||
8. **Provide incident response** procedures and escalation workflows
|
||||
|
||||
## Example Interactions
|
||||
- "Design a comprehensive monitoring strategy for a microservices architecture with 50+ services"
|
||||
- "Implement distributed tracing for a complex e-commerce platform handling 1M+ daily transactions"
|
||||
- "Set up cost-effective log management for a high-traffic application generating 10TB+ daily logs"
|
||||
- "Create SLI/SLO framework with error budget tracking for API services with 99.9% availability target"
|
||||
- "Build real-time alerting system with intelligent noise reduction for 24/7 operations team"
|
||||
- "Implement chaos engineering with monitoring validation for Netflix-scale resilience testing"
|
||||
- "Design executive dashboard showing business impact of system reliability and revenue correlation"
|
||||
- "Set up compliance monitoring for SOC2 and PCI requirements with automated evidence collection"
|
||||
- "Optimize monitoring costs while maintaining comprehensive coverage for startup scaling to enterprise"
|
||||
- "Create automated incident response workflows with runbook integration and Slack/PagerDuty escalation"
|
||||
- "Build multi-region observability architecture with data sovereignty compliance"
|
||||
- "Implement machine learning-based anomaly detection for proactive issue identification"
|
||||
- "Design observability strategy for serverless architecture with AWS Lambda and API Gateway"
|
||||
- "Create custom metrics pipeline for business KPIs integrated with technical monitoring"
|
||||
|
||||
## Limitations
|
||||
- Use this skill only when the task clearly matches the scope described above.
|
||||
- Do not treat the output as a substitute for environment-specific validation, testing, or expert review.
|
||||
- Stop and ask for clarification if required inputs, permissions, safety boundaries, or success criteria are missing.
|
||||
+181
@@ -0,0 +1,181 @@
|
||||
---
|
||||
name: performance-engineer
|
||||
description: "Expert performance engineer specializing in modern observability,"
|
||||
risk: unknown
|
||||
source: community
|
||||
date_added: "2026-02-27"
|
||||
---
|
||||
You are a performance engineer specializing in modern application optimization, observability, and scalable system performance.
|
||||
|
||||
## Use this skill when
|
||||
|
||||
- Diagnosing performance bottlenecks in backend, frontend, or infrastructure
|
||||
- Designing load tests, capacity plans, or scalability strategies
|
||||
- Setting up observability and performance monitoring
|
||||
- Optimizing latency, throughput, or resource efficiency
|
||||
|
||||
## Do not use this skill when
|
||||
|
||||
- The task is feature development with no performance goals
|
||||
- There is no access to metrics, traces, or profiling data
|
||||
- A quick, non-technical summary is the only requirement
|
||||
|
||||
## Instructions
|
||||
|
||||
1. Confirm performance goals, user impact, and baseline metrics.
|
||||
2. Collect traces, profiles, and load tests to isolate bottlenecks.
|
||||
3. Propose optimizations with expected impact and tradeoffs.
|
||||
4. Verify results and add guardrails to prevent regressions.
|
||||
|
||||
## Safety
|
||||
|
||||
- Avoid load testing production without approvals and safeguards.
|
||||
- Use staged rollouts with rollback plans for high-risk changes.
|
||||
|
||||
## Purpose
|
||||
Expert performance engineer with comprehensive knowledge of modern observability, application profiling, and system optimization. Masters performance testing, distributed tracing, caching architectures, and scalability patterns. Specializes in end-to-end performance optimization, real user monitoring, and building performant, scalable systems.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### Modern Observability & Monitoring
|
||||
- **OpenTelemetry**: Distributed tracing, metrics collection, correlation across services
|
||||
- **APM platforms**: DataDog APM, New Relic, Dynatrace, AppDynamics, Honeycomb, Jaeger
|
||||
- **Metrics & monitoring**: Prometheus, Grafana, InfluxDB, custom metrics, SLI/SLO tracking
|
||||
- **Real User Monitoring (RUM)**: User experience tracking, Core Web Vitals, page load analytics
|
||||
- **Synthetic monitoring**: Uptime monitoring, API testing, user journey simulation
|
||||
- **Log correlation**: Structured logging, distributed log tracing, error correlation
|
||||
|
||||
### Advanced Application Profiling
|
||||
- **CPU profiling**: Flame graphs, call stack analysis, hotspot identification
|
||||
- **Memory profiling**: Heap analysis, garbage collection tuning, memory leak detection
|
||||
- **I/O profiling**: Disk I/O optimization, network latency analysis, database query profiling
|
||||
- **Language-specific profiling**: JVM profiling, Python profiling, Node.js profiling, Go profiling
|
||||
- **Container profiling**: Docker performance analysis, Kubernetes resource optimization
|
||||
- **Cloud profiling**: AWS X-Ray, Azure Application Insights, GCP Cloud Profiler
|
||||
|
||||
### Modern Load Testing & Performance Validation
|
||||
- **Load testing tools**: k6, JMeter, Gatling, Locust, Artillery, cloud-based testing
|
||||
- **API testing**: REST API testing, GraphQL performance testing, WebSocket testing
|
||||
- **Browser testing**: Puppeteer, Playwright, Selenium WebDriver performance testing
|
||||
- **Chaos engineering**: Netflix Chaos Monkey, Gremlin, failure injection testing
|
||||
- **Performance budgets**: Budget tracking, CI/CD integration, regression detection
|
||||
- **Scalability testing**: Auto-scaling validation, capacity planning, breaking point analysis
|
||||
|
||||
### Multi-Tier Caching Strategies
|
||||
- **Application caching**: In-memory caching, object caching, computed value caching
|
||||
- **Distributed caching**: Redis, Memcached, Hazelcast, cloud cache services
|
||||
- **Database caching**: Query result caching, connection pooling, buffer pool optimization
|
||||
- **CDN optimization**: CloudFlare, AWS CloudFront, Azure CDN, edge caching strategies
|
||||
- **Browser caching**: HTTP cache headers, service workers, offline-first strategies
|
||||
- **API caching**: Response caching, conditional requests, cache invalidation strategies
|
||||
|
||||
### Frontend Performance Optimization
|
||||
- **Core Web Vitals**: LCP, FID, CLS optimization, Web Performance API
|
||||
- **Resource optimization**: Image optimization, lazy loading, critical resource prioritization
|
||||
- **JavaScript optimization**: Bundle splitting, tree shaking, code splitting, lazy loading
|
||||
- **CSS optimization**: Critical CSS, CSS optimization, render-blocking resource elimination
|
||||
- **Network optimization**: HTTP/2, HTTP/3, resource hints, preloading strategies
|
||||
- **Progressive Web Apps**: Service workers, caching strategies, offline functionality
|
||||
|
||||
### Backend Performance Optimization
|
||||
- **API optimization**: Response time optimization, pagination, bulk operations
|
||||
- **Microservices performance**: Service-to-service optimization, circuit breakers, bulkheads
|
||||
- **Async processing**: Background jobs, message queues, event-driven architectures
|
||||
- **Database optimization**: Query optimization, indexing, connection pooling, read replicas
|
||||
- **Concurrency optimization**: Thread pool tuning, async/await patterns, resource locking
|
||||
- **Resource management**: CPU optimization, memory management, garbage collection tuning
|
||||
|
||||
### Distributed System Performance
|
||||
- **Service mesh optimization**: Istio, Linkerd performance tuning, traffic management
|
||||
- **Message queue optimization**: Kafka, RabbitMQ, SQS performance tuning
|
||||
- **Event streaming**: Real-time processing optimization, stream processing performance
|
||||
- **API gateway optimization**: Rate limiting, caching, traffic shaping
|
||||
- **Load balancing**: Traffic distribution, health checks, failover optimization
|
||||
- **Cross-service communication**: gRPC optimization, REST API performance, GraphQL optimization
|
||||
|
||||
### Cloud Performance Optimization
|
||||
- **Auto-scaling optimization**: HPA, VPA, cluster autoscaling, scaling policies
|
||||
- **Serverless optimization**: Lambda performance, cold start optimization, memory allocation
|
||||
- **Container optimization**: Docker image optimization, Kubernetes resource limits
|
||||
- **Network optimization**: VPC performance, CDN integration, edge computing
|
||||
- **Storage optimization**: Disk I/O performance, database performance, object storage
|
||||
- **Cost-performance optimization**: Right-sizing, reserved capacity, spot instances
|
||||
|
||||
### Performance Testing Automation
|
||||
- **CI/CD integration**: Automated performance testing, regression detection
|
||||
- **Performance gates**: Automated pass/fail criteria, deployment blocking
|
||||
- **Continuous profiling**: Production profiling, performance trend analysis
|
||||
- **A/B testing**: Performance comparison, canary analysis, feature flag performance
|
||||
- **Regression testing**: Automated performance regression detection, baseline management
|
||||
- **Capacity testing**: Load testing automation, capacity planning validation
|
||||
|
||||
### Database & Data Performance
|
||||
- **Query optimization**: Execution plan analysis, index optimization, query rewriting
|
||||
- **Connection optimization**: Connection pooling, prepared statements, batch processing
|
||||
- **Caching strategies**: Query result caching, object-relational mapping optimization
|
||||
- **Data pipeline optimization**: ETL performance, streaming data processing
|
||||
- **NoSQL optimization**: MongoDB, DynamoDB, Redis performance tuning
|
||||
- **Time-series optimization**: InfluxDB, TimescaleDB, metrics storage optimization
|
||||
|
||||
### Mobile & Edge Performance
|
||||
- **Mobile optimization**: React Native, Flutter performance, native app optimization
|
||||
- **Edge computing**: CDN performance, edge functions, geo-distributed optimization
|
||||
- **Network optimization**: Mobile network performance, offline-first strategies
|
||||
- **Battery optimization**: CPU usage optimization, background processing efficiency
|
||||
- **User experience**: Touch responsiveness, smooth animations, perceived performance
|
||||
|
||||
### Performance Analytics & Insights
|
||||
- **User experience analytics**: Session replay, heatmaps, user behavior analysis
|
||||
- **Performance budgets**: Resource budgets, timing budgets, metric tracking
|
||||
- **Business impact analysis**: Performance-revenue correlation, conversion optimization
|
||||
- **Competitive analysis**: Performance benchmarking, industry comparison
|
||||
- **ROI analysis**: Performance optimization impact, cost-benefit analysis
|
||||
- **Alerting strategies**: Performance anomaly detection, proactive alerting
|
||||
|
||||
## Behavioral Traits
|
||||
- Measures performance comprehensively before implementing any optimizations
|
||||
- Focuses on the biggest bottlenecks first for maximum impact and ROI
|
||||
- Sets and enforces performance budgets to prevent regression
|
||||
- Implements caching at appropriate layers with proper invalidation strategies
|
||||
- Conducts load testing with realistic scenarios and production-like data
|
||||
- Prioritizes user-perceived performance over synthetic benchmarks
|
||||
- Uses data-driven decision making with comprehensive metrics and monitoring
|
||||
- Considers the entire system architecture when optimizing performance
|
||||
- Balances performance optimization with maintainability and cost
|
||||
- Implements continuous performance monitoring and alerting
|
||||
|
||||
## Knowledge Base
|
||||
- Modern observability platforms and distributed tracing technologies
|
||||
- Application profiling tools and performance analysis methodologies
|
||||
- Load testing strategies and performance validation techniques
|
||||
- Caching architectures and strategies across different system layers
|
||||
- Frontend and backend performance optimization best practices
|
||||
- Cloud platform performance characteristics and optimization opportunities
|
||||
- Database performance tuning and optimization techniques
|
||||
- Distributed system performance patterns and anti-patterns
|
||||
|
||||
## Response Approach
|
||||
1. **Establish performance baseline** with comprehensive measurement and profiling
|
||||
2. **Identify critical bottlenecks** through systematic analysis and user journey mapping
|
||||
3. **Prioritize optimizations** based on user impact, business value, and implementation effort
|
||||
4. **Implement optimizations** with proper testing and validation procedures
|
||||
5. **Set up monitoring and alerting** for continuous performance tracking
|
||||
6. **Validate improvements** through comprehensive testing and user experience measurement
|
||||
7. **Establish performance budgets** to prevent future regression
|
||||
8. **Document optimizations** with clear metrics and impact analysis
|
||||
9. **Plan for scalability** with appropriate caching and architectural improvements
|
||||
|
||||
## Example Interactions
|
||||
- "Analyze and optimize end-to-end API performance with distributed tracing and caching"
|
||||
- "Implement comprehensive observability stack with OpenTelemetry, Prometheus, and Grafana"
|
||||
- "Optimize React application for Core Web Vitals and user experience metrics"
|
||||
- "Design load testing strategy for microservices architecture with realistic traffic patterns"
|
||||
- "Implement multi-tier caching architecture for high-traffic e-commerce application"
|
||||
- "Optimize database performance for analytical workloads with query and index optimization"
|
||||
- "Create performance monitoring dashboard with SLI/SLO tracking and automated alerting"
|
||||
- "Implement chaos engineering practices for distributed system resilience and performance validation"
|
||||
|
||||
## Limitations
|
||||
- Use this skill only when the task clearly matches the scope described above.
|
||||
- Do not treat the output as a substitute for environment-specific validation, testing, or expert review.
|
||||
- Stop and ask for clarification if required inputs, permissions, safety boundaries, or success criteria are missing.
|
||||
+394
@@ -0,0 +1,394 @@
|
||||
---
|
||||
name: postmortem-writing
|
||||
description: "Comprehensive guide to writing effective, blameless postmortems that drive organizational learning and prevent incident recurrence."
|
||||
risk: unknown
|
||||
source: community
|
||||
date_added: "2026-02-27"
|
||||
---
|
||||
|
||||
# Postmortem Writing
|
||||
|
||||
Comprehensive guide to writing effective, blameless postmortems that drive organizational learning and prevent incident recurrence.
|
||||
|
||||
## Do not use this skill when
|
||||
|
||||
- The task is unrelated to postmortem writing
|
||||
- You need a different domain or tool outside this scope
|
||||
|
||||
## Instructions
|
||||
|
||||
- Clarify goals, constraints, and required inputs.
|
||||
- Apply relevant best practices and validate outcomes.
|
||||
- Provide actionable steps and verification.
|
||||
- If detailed examples are required, open `resources/implementation-playbook.md`.
|
||||
|
||||
## Use this skill when
|
||||
|
||||
- Conducting post-incident reviews
|
||||
- Writing postmortem documents
|
||||
- Facilitating blameless postmortem meetings
|
||||
- Identifying root causes and contributing factors
|
||||
- Creating actionable follow-up items
|
||||
- Building organizational learning culture
|
||||
|
||||
## Core Concepts
|
||||
|
||||
### 1. Blameless Culture
|
||||
|
||||
| Blame-Focused | Blameless |
|
||||
|---------------|-----------|
|
||||
| "Who caused this?" | "What conditions allowed this?" |
|
||||
| "Someone made a mistake" | "The system allowed this mistake" |
|
||||
| Punish individuals | Improve systems |
|
||||
| Hide information | Share learnings |
|
||||
| Fear of speaking up | Psychological safety |
|
||||
|
||||
### 2. Postmortem Triggers
|
||||
|
||||
- SEV1 or SEV2 incidents
|
||||
- Customer-facing outages > 15 minutes
|
||||
- Data loss or security incidents
|
||||
- Near-misses that could have been severe
|
||||
- Novel failure modes
|
||||
- Incidents requiring unusual intervention
|
||||
|
||||
## Quick Start
|
||||
|
||||
### Postmortem Timeline
|
||||
```
|
||||
Day 0: Incident occurs
|
||||
Day 1-2: Draft postmortem document
|
||||
Day 3-5: Postmortem meeting
|
||||
Day 5-7: Finalize document, create tickets
|
||||
Week 2+: Action item completion
|
||||
Quarterly: Review patterns across incidents
|
||||
```
|
||||
|
||||
## Templates
|
||||
|
||||
### Template 1: Standard Postmortem
|
||||
|
||||
```markdown
|
||||
# Postmortem: [Incident Title]
|
||||
|
||||
**Date**: 2024-01-15
|
||||
**Authors**: @alice, @bob
|
||||
**Status**: Draft | In Review | Final
|
||||
**Incident Severity**: SEV2
|
||||
**Incident Duration**: 47 minutes
|
||||
|
||||
## Executive Summary
|
||||
|
||||
On January 15, 2024, the payment processing service experienced a 47-minute outage affecting approximately 12,000 customers. The root cause was a database connection pool exhaustion triggered by a configuration change in deployment v2.3.4. The incident was resolved by rolling back to v2.3.3 and increasing connection pool limits.
|
||||
|
||||
**Impact**:
|
||||
- 12,000 customers unable to complete purchases
|
||||
- Estimated revenue loss: $45,000
|
||||
- 847 support tickets created
|
||||
- No data loss or security implications
|
||||
|
||||
## Timeline (All times UTC)
|
||||
|
||||
| Time | Event |
|
||||
|------|-------|
|
||||
| 14:23 | Deployment v2.3.4 completed to production |
|
||||
| 14:31 | First alert: `payment_error_rate > 5%` |
|
||||
| 14:33 | On-call engineer @alice acknowledges alert |
|
||||
| 14:35 | Initial investigation begins, error rate at 23% |
|
||||
| 14:41 | Incident declared SEV2, @bob joins |
|
||||
| 14:45 | Database connection exhaustion identified |
|
||||
| 14:52 | Decision to rollback deployment |
|
||||
| 14:58 | Rollback to v2.3.3 initiated |
|
||||
| 15:10 | Rollback complete, error rate dropping |
|
||||
| 15:18 | Service fully recovered, incident resolved |
|
||||
|
||||
## Root Cause Analysis
|
||||
|
||||
### What Happened
|
||||
|
||||
The v2.3.4 deployment included a change to the database query pattern that inadvertently removed connection pooling for a frequently-called endpoint. Each request opened a new database connection instead of reusing pooled connections.
|
||||
|
||||
### Why It Happened
|
||||
|
||||
1. **Proximate Cause**: Code change in `PaymentRepository.java` replaced pooled `DataSource` with direct `DriverManager.getConnection()` calls.
|
||||
|
||||
2. **Contributing Factors**:
|
||||
- Code review did not catch the connection handling change
|
||||
- No integration tests specifically for connection pool behavior
|
||||
- Staging environment has lower traffic, masking the issue
|
||||
- Database connection metrics alert threshold was too high (90%)
|
||||
|
||||
3. **5 Whys Analysis**:
|
||||
- Why did the service fail? → Database connections exhausted
|
||||
- Why were connections exhausted? → Each request opened new connection
|
||||
- Why did each request open new connection? → Code bypassed connection pool
|
||||
- Why did code bypass connection pool? → Developer unfamiliar with codebase patterns
|
||||
- Why was developer unfamiliar? → No documentation on connection management patterns
|
||||
|
||||
### System Diagram
|
||||
|
||||
```
|
||||
[Client] → [Load Balancer] → [Payment Service] → [Database]
|
||||
↓
|
||||
Connection Pool (broken)
|
||||
↓
|
||||
Direct connections (cause)
|
||||
```
|
||||
|
||||
## Detection
|
||||
|
||||
### What Worked
|
||||
- Error rate alert fired within 8 minutes of deployment
|
||||
- Grafana dashboard clearly showed connection spike
|
||||
- On-call response was swift (2 minute acknowledgment)
|
||||
|
||||
### What Didn't Work
|
||||
- Database connection metric alert threshold too high
|
||||
- No deployment-correlated alerting
|
||||
- Canary deployment would have caught this earlier
|
||||
|
||||
### Detection Gap
|
||||
The deployment completed at 14:23, but the first alert didn't fire until 14:31 (8 minutes). A deployment-aware alert could have detected the issue faster.
|
||||
|
||||
## Response
|
||||
|
||||
### What Worked
|
||||
- On-call engineer quickly identified database as the issue
|
||||
- Rollback decision was made decisively
|
||||
- Clear communication in incident channel
|
||||
|
||||
### What Could Be Improved
|
||||
- Took 10 minutes to correlate issue with recent deployment
|
||||
- Had to manually check deployment history
|
||||
- Rollback took 12 minutes (could be faster)
|
||||
|
||||
## Impact
|
||||
|
||||
### Customer Impact
|
||||
- 12,000 unique customers affected
|
||||
- Average impact duration: 35 minutes
|
||||
- 847 support tickets (23% of affected users)
|
||||
- Customer satisfaction score dropped 12 points
|
||||
|
||||
### Business Impact
|
||||
- Estimated revenue loss: $45,000
|
||||
- Support cost: ~$2,500 (agent time)
|
||||
- Engineering time: ~8 person-hours
|
||||
|
||||
### Technical Impact
|
||||
- Database primary experienced elevated load
|
||||
- Some replica lag during incident
|
||||
- No permanent damage to systems
|
||||
|
||||
## Lessons Learned
|
||||
|
||||
### What Went Well
|
||||
1. Alerting detected the issue before customer reports
|
||||
2. Team collaborated effectively under pressure
|
||||
3. Rollback procedure worked smoothly
|
||||
4. Communication was clear and timely
|
||||
|
||||
### What Went Wrong
|
||||
1. Code review missed critical change
|
||||
2. Test coverage gap for connection pooling
|
||||
3. Staging environment doesn't reflect production traffic
|
||||
4. Alert thresholds were not tuned properly
|
||||
|
||||
### Where We Got Lucky
|
||||
1. Incident occurred during business hours with full team available
|
||||
2. Database handled the load without failing completely
|
||||
3. No other incidents occurred simultaneously
|
||||
|
||||
## Action Items
|
||||
|
||||
| Priority | Action | Owner | Due Date | Ticket |
|
||||
|----------|--------|-------|----------|--------|
|
||||
| P0 | Add integration test for connection pool behavior | @alice | 2024-01-22 | ENG-1234 |
|
||||
| P0 | Lower database connection alert threshold to 70% | @bob | 2024-01-17 | OPS-567 |
|
||||
| P1 | Document connection management patterns | @alice | 2024-01-29 | DOC-89 |
|
||||
| P1 | Implement deployment-correlated alerting | @bob | 2024-02-05 | OPS-568 |
|
||||
| P2 | Evaluate canary deployment strategy | @charlie | 2024-02-15 | ENG-1235 |
|
||||
| P2 | Load test staging with production-like traffic | @dave | 2024-02-28 | QA-123 |
|
||||
|
||||
## Appendix
|
||||
|
||||
### Supporting Data
|
||||
|
||||
#### Error Rate Graph
|
||||
[Link to Grafana dashboard snapshot]
|
||||
|
||||
#### Database Connection Graph
|
||||
[Link to metrics]
|
||||
|
||||
### Related Incidents
|
||||
- 2023-11-02: Similar connection issue in User Service (POSTMORTEM-42)
|
||||
|
||||
### References
|
||||
- Connection Pool Best Practices
|
||||
- Deployment Runbook
|
||||
```
|
||||
|
||||
### Template 2: 5 Whys Analysis
|
||||
|
||||
```markdown
|
||||
# 5 Whys Analysis: [Incident]
|
||||
|
||||
## Problem Statement
|
||||
Payment service experienced 47-minute outage due to database connection exhaustion.
|
||||
|
||||
## Analysis
|
||||
|
||||
### Why #1: Why did the service fail?
|
||||
**Answer**: Database connections were exhausted, causing all new requests to fail.
|
||||
|
||||
**Evidence**: Metrics showed connection count at 100/100 (max), with 500+ pending requests.
|
||||
|
||||
---
|
||||
|
||||
### Why #2: Why were database connections exhausted?
|
||||
**Answer**: Each incoming request opened a new database connection instead of using the connection pool.
|
||||
|
||||
**Evidence**: Code diff shows direct `DriverManager.getConnection()` instead of pooled `DataSource`.
|
||||
|
||||
---
|
||||
|
||||
### Why #3: Why did the code bypass the connection pool?
|
||||
**Answer**: A developer refactored the repository class and inadvertently changed the connection acquisition method.
|
||||
|
||||
**Evidence**: PR #1234 shows the change, made while fixing a different bug.
|
||||
|
||||
---
|
||||
|
||||
### Why #4: Why wasn't this caught in code review?
|
||||
**Answer**: The reviewer focused on the functional change (the bug fix) and didn't notice the infrastructure change.
|
||||
|
||||
**Evidence**: Review comments only discuss business logic.
|
||||
|
||||
---
|
||||
|
||||
### Why #5: Why isn't there a safety net for this type of change?
|
||||
**Answer**: We lack automated tests that verify connection pool behavior and lack documentation about our connection patterns.
|
||||
|
||||
**Evidence**: Test suite has no tests for connection handling; wiki has no article on database connections.
|
||||
|
||||
## Root Causes Identified
|
||||
|
||||
1. **Primary**: Missing automated tests for infrastructure behavior
|
||||
2. **Secondary**: Insufficient documentation of architectural patterns
|
||||
3. **Tertiary**: Code review checklist doesn't include infrastructure considerations
|
||||
|
||||
## Systemic Improvements
|
||||
|
||||
| Root Cause | Improvement | Type |
|
||||
|------------|-------------|------|
|
||||
| Missing tests | Add infrastructure behavior tests | Prevention |
|
||||
| Missing docs | Document connection patterns | Prevention |
|
||||
| Review gaps | Update review checklist | Detection |
|
||||
| No canary | Implement canary deployments | Mitigation |
|
||||
```
|
||||
|
||||
### Template 3: Quick Postmortem (Minor Incidents)
|
||||
|
||||
```markdown
|
||||
# Quick Postmortem: [Brief Title]
|
||||
|
||||
**Date**: 2024-01-15 | **Duration**: 12 min | **Severity**: SEV3
|
||||
|
||||
## What Happened
|
||||
API latency spiked to 5s due to cache miss storm after cache flush.
|
||||
|
||||
## Timeline
|
||||
- 10:00 - Cache flush initiated for config update
|
||||
- 10:02 - Latency alerts fire
|
||||
- 10:05 - Identified as cache miss storm
|
||||
- 10:08 - Enabled cache warming
|
||||
- 10:12 - Latency normalized
|
||||
|
||||
## Root Cause
|
||||
Full cache flush for minor config update caused thundering herd.
|
||||
|
||||
## Fix
|
||||
- Immediate: Enabled cache warming
|
||||
- Long-term: Implement partial cache invalidation (ENG-999)
|
||||
|
||||
## Lessons
|
||||
Don't full-flush cache in production; use targeted invalidation.
|
||||
```
|
||||
|
||||
## Facilitation Guide
|
||||
|
||||
### Running a Postmortem Meeting
|
||||
|
||||
```markdown
|
||||
## Meeting Structure (60 minutes)
|
||||
|
||||
### 1. Opening (5 min)
|
||||
- Remind everyone of blameless culture
|
||||
- "We're here to learn, not to blame"
|
||||
- Review meeting norms
|
||||
|
||||
### 2. Timeline Review (15 min)
|
||||
- Walk through events chronologically
|
||||
- Ask clarifying questions
|
||||
- Identify gaps in timeline
|
||||
|
||||
### 3. Analysis Discussion (20 min)
|
||||
- What failed?
|
||||
- Why did it fail?
|
||||
- What conditions allowed this?
|
||||
- What would have prevented it?
|
||||
|
||||
### 4. Action Items (15 min)
|
||||
- Brainstorm improvements
|
||||
- Prioritize by impact and effort
|
||||
- Assign owners and due dates
|
||||
|
||||
### 5. Closing (5 min)
|
||||
- Summarize key learnings
|
||||
- Confirm action item owners
|
||||
- Schedule follow-up if needed
|
||||
|
||||
## Facilitation Tips
|
||||
- Keep discussion on track
|
||||
- Redirect blame to systems
|
||||
- Encourage quiet participants
|
||||
- Document dissenting views
|
||||
- Time-box tangents
|
||||
```
|
||||
|
||||
## Anti-Patterns to Avoid
|
||||
|
||||
| Anti-Pattern | Problem | Better Approach |
|
||||
|--------------|---------|-----------------|
|
||||
| **Blame game** | Shuts down learning | Focus on systems |
|
||||
| **Shallow analysis** | Doesn't prevent recurrence | Ask "why" 5 times |
|
||||
| **No action items** | Waste of time | Always have concrete next steps |
|
||||
| **Unrealistic actions** | Never completed | Scope to achievable tasks |
|
||||
| **No follow-up** | Actions forgotten | Track in ticketing system |
|
||||
|
||||
## Best Practices
|
||||
|
||||
### Do's
|
||||
- **Start immediately** - Memory fades fast
|
||||
- **Be specific** - Exact times, exact errors
|
||||
- **Include graphs** - Visual evidence
|
||||
- **Assign owners** - No orphan action items
|
||||
- **Share widely** - Organizational learning
|
||||
|
||||
### Don'ts
|
||||
- **Don't name and shame** - Ever
|
||||
- **Don't skip small incidents** - They reveal patterns
|
||||
- **Don't make it a blame doc** - That kills learning
|
||||
- **Don't create busywork** - Actions should be meaningful
|
||||
- **Don't skip follow-up** - Verify actions completed
|
||||
|
||||
## Resources
|
||||
|
||||
- [Google SRE - Postmortem Culture](https://sre.google/sre-book/postmortem-culture/)
|
||||
- [Etsy's Blameless Postmortems](https://codeascraft.com/2012/05/22/blameless-postmortems/)
|
||||
- [PagerDuty Postmortem Guide](https://postmortems.pagerduty.com/)
|
||||
|
||||
## Limitations
|
||||
- Use this skill only when the task clearly matches the scope described above.
|
||||
- Do not treat the output as a substitute for environment-specific validation, testing, or expert review.
|
||||
- Stop and ask for clarification if required inputs, permissions, safety boundaries, or success criteria are missing.
|
||||
+349
@@ -0,0 +1,349 @@
|
||||
---
|
||||
name: slo-implementation
|
||||
description: "Framework for defining and implementing Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets."
|
||||
risk: unknown
|
||||
source: community
|
||||
date_added: "2026-02-27"
|
||||
---
|
||||
|
||||
# SLO Implementation
|
||||
|
||||
Framework for defining and implementing Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
|
||||
|
||||
## Do not use this skill when
|
||||
|
||||
- The task is unrelated to slo implementation
|
||||
- You need a different domain or tool outside this scope
|
||||
|
||||
## Instructions
|
||||
|
||||
- Clarify goals, constraints, and required inputs.
|
||||
- Apply relevant best practices and validate outcomes.
|
||||
- Provide actionable steps and verification.
|
||||
- If detailed examples are required, open `resources/implementation-playbook.md`.
|
||||
|
||||
## Purpose
|
||||
|
||||
Implement measurable reliability targets using SLIs, SLOs, and error budgets to balance reliability with innovation velocity.
|
||||
|
||||
## Use this skill when
|
||||
|
||||
- Define service reliability targets
|
||||
- Measure user-perceived reliability
|
||||
- Implement error budgets
|
||||
- Create SLO-based alerts
|
||||
- Track reliability goals
|
||||
|
||||
## SLI/SLO/SLA Hierarchy
|
||||
|
||||
```
|
||||
SLA (Service Level Agreement)
|
||||
↓ Contract with customers
|
||||
SLO (Service Level Objective)
|
||||
↓ Internal reliability target
|
||||
SLI (Service Level Indicator)
|
||||
↓ Actual measurement
|
||||
```
|
||||
|
||||
## Defining SLIs
|
||||
|
||||
### Common SLI Types
|
||||
|
||||
#### 1. Availability SLI
|
||||
```promql
|
||||
# Successful requests / Total requests
|
||||
sum(rate(http_requests_total{status!~"5.."}[28d]))
|
||||
/
|
||||
sum(rate(http_requests_total[28d]))
|
||||
```
|
||||
|
||||
#### 2. Latency SLI
|
||||
```promql
|
||||
# Requests below latency threshold / Total requests
|
||||
sum(rate(http_request_duration_seconds_bucket{le="0.5"}[28d]))
|
||||
/
|
||||
sum(rate(http_request_duration_seconds_count[28d]))
|
||||
```
|
||||
|
||||
#### 3. Durability SLI
|
||||
```
|
||||
# Successful writes / Total writes
|
||||
sum(storage_writes_successful_total)
|
||||
/
|
||||
sum(storage_writes_total)
|
||||
```
|
||||
|
||||
**Reference:** See `references/slo-definitions.md`
|
||||
|
||||
## Setting SLO Targets
|
||||
|
||||
### Availability SLO Examples
|
||||
|
||||
| SLO % | Downtime/Month | Downtime/Year |
|
||||
|-------|----------------|---------------|
|
||||
| 99% | 7.2 hours | 3.65 days |
|
||||
| 99.9% | 43.2 minutes | 8.76 hours |
|
||||
| 99.95%| 21.6 minutes | 4.38 hours |
|
||||
| 99.99%| 4.32 minutes | 52.56 minutes |
|
||||
|
||||
### Choose Appropriate SLOs
|
||||
|
||||
**Consider:**
|
||||
- User expectations
|
||||
- Business requirements
|
||||
- Current performance
|
||||
- Cost of reliability
|
||||
- Competitor benchmarks
|
||||
|
||||
**Example SLOs:**
|
||||
```yaml
|
||||
slos:
|
||||
- name: api_availability
|
||||
target: 99.9
|
||||
window: 28d
|
||||
sli: |
|
||||
sum(rate(http_requests_total{status!~"5.."}[28d]))
|
||||
/
|
||||
sum(rate(http_requests_total[28d]))
|
||||
|
||||
- name: api_latency_p95
|
||||
target: 99
|
||||
window: 28d
|
||||
sli: |
|
||||
sum(rate(http_request_duration_seconds_bucket{le="0.5"}[28d]))
|
||||
/
|
||||
sum(rate(http_request_duration_seconds_count[28d]))
|
||||
```
|
||||
|
||||
## Error Budget Calculation
|
||||
|
||||
### Error Budget Formula
|
||||
|
||||
```
|
||||
Error Budget = 1 - SLO Target
|
||||
```
|
||||
|
||||
**Example:**
|
||||
- SLO: 99.9% availability
|
||||
- Error Budget: 0.1% = 43.2 minutes/month
|
||||
- Current Error: 0.05% = 21.6 minutes/month
|
||||
- Remaining Budget: 50%
|
||||
|
||||
### Error Budget Policy
|
||||
|
||||
```yaml
|
||||
error_budget_policy:
|
||||
- remaining_budget: 100%
|
||||
action: Normal development velocity
|
||||
- remaining_budget: 50%
|
||||
action: Consider postponing risky changes
|
||||
- remaining_budget: 10%
|
||||
action: Freeze non-critical changes
|
||||
- remaining_budget: 0%
|
||||
action: Feature freeze, focus on reliability
|
||||
```
|
||||
|
||||
**Reference:** See `references/error-budget.md`
|
||||
|
||||
## SLO Implementation
|
||||
|
||||
### Prometheus Recording Rules
|
||||
|
||||
```yaml
|
||||
# SLI Recording Rules
|
||||
groups:
|
||||
- name: sli_rules
|
||||
interval: 30s
|
||||
rules:
|
||||
# Availability SLI
|
||||
- record: sli:http_availability:ratio
|
||||
expr: |
|
||||
sum(rate(http_requests_total{status!~"5.."}[28d]))
|
||||
/
|
||||
sum(rate(http_requests_total[28d]))
|
||||
|
||||
# Latency SLI (requests < 500ms)
|
||||
- record: sli:http_latency:ratio
|
||||
expr: |
|
||||
sum(rate(http_request_duration_seconds_bucket{le="0.5"}[28d]))
|
||||
/
|
||||
sum(rate(http_request_duration_seconds_count[28d]))
|
||||
|
||||
- name: slo_rules
|
||||
interval: 5m
|
||||
rules:
|
||||
# SLO compliance (1 = meeting SLO, 0 = violating)
|
||||
- record: slo:http_availability:compliance
|
||||
expr: sli:http_availability:ratio >= bool 0.999
|
||||
|
||||
- record: slo:http_latency:compliance
|
||||
expr: sli:http_latency:ratio >= bool 0.99
|
||||
|
||||
# Error budget remaining (percentage)
|
||||
- record: slo:http_availability:error_budget_remaining
|
||||
expr: |
|
||||
(sli:http_availability:ratio - 0.999) / (1 - 0.999) * 100
|
||||
|
||||
# Error budget burn rate
|
||||
- record: slo:http_availability:burn_rate_5m
|
||||
expr: |
|
||||
(1 - (
|
||||
sum(rate(http_requests_total{status!~"5.."}[5m]))
|
||||
/
|
||||
sum(rate(http_requests_total[5m]))
|
||||
)) / (1 - 0.999)
|
||||
```
|
||||
|
||||
### SLO Alerting Rules
|
||||
|
||||
```yaml
|
||||
groups:
|
||||
- name: slo_alerts
|
||||
interval: 1m
|
||||
rules:
|
||||
# Fast burn: 14.4x rate, 1 hour window
|
||||
# Consumes 2% error budget in 1 hour
|
||||
- alert: SLOErrorBudgetBurnFast
|
||||
expr: |
|
||||
slo:http_availability:burn_rate_1h > 14.4
|
||||
and
|
||||
slo:http_availability:burn_rate_5m > 14.4
|
||||
for: 2m
|
||||
labels:
|
||||
severity: critical
|
||||
annotations:
|
||||
summary: "Fast error budget burn detected"
|
||||
description: "Error budget burning at {{ $value }}x rate"
|
||||
|
||||
# Slow burn: 6x rate, 6 hour window
|
||||
# Consumes 5% error budget in 6 hours
|
||||
- alert: SLOErrorBudgetBurnSlow
|
||||
expr: |
|
||||
slo:http_availability:burn_rate_6h > 6
|
||||
and
|
||||
slo:http_availability:burn_rate_30m > 6
|
||||
for: 15m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "Slow error budget burn detected"
|
||||
description: "Error budget burning at {{ $value }}x rate"
|
||||
|
||||
# Error budget exhausted
|
||||
- alert: SLOErrorBudgetExhausted
|
||||
expr: slo:http_availability:error_budget_remaining < 0
|
||||
for: 5m
|
||||
labels:
|
||||
severity: critical
|
||||
annotations:
|
||||
summary: "SLO error budget exhausted"
|
||||
description: "Error budget remaining: {{ $value }}%"
|
||||
```
|
||||
|
||||
## SLO Dashboard
|
||||
|
||||
**Grafana Dashboard Structure:**
|
||||
|
||||
```
|
||||
┌────────────────────────────────────┐
|
||||
│ SLO Compliance (Current) │
|
||||
│ ✓ 99.95% (Target: 99.9%) │
|
||||
├────────────────────────────────────┤
|
||||
│ Error Budget Remaining: 65% │
|
||||
│ ████████░░ 65% │
|
||||
├────────────────────────────────────┤
|
||||
│ SLI Trend (28 days) │
|
||||
│ [Time series graph] │
|
||||
├────────────────────────────────────┤
|
||||
│ Burn Rate Analysis │
|
||||
│ [Burn rate by time window] │
|
||||
└────────────────────────────────────┘
|
||||
```
|
||||
|
||||
**Example Queries:**
|
||||
|
||||
```promql
|
||||
# Current SLO compliance
|
||||
sli:http_availability:ratio * 100
|
||||
|
||||
# Error budget remaining
|
||||
slo:http_availability:error_budget_remaining
|
||||
|
||||
# Days until error budget exhausted (at current burn rate)
|
||||
(slo:http_availability:error_budget_remaining / 100)
|
||||
*
|
||||
28
|
||||
/
|
||||
(1 - sli:http_availability:ratio) * (1 - 0.999)
|
||||
```
|
||||
|
||||
## Multi-Window Burn Rate Alerts
|
||||
|
||||
```yaml
|
||||
# Combination of short and long windows reduces false positives
|
||||
rules:
|
||||
- alert: SLOBurnRateHigh
|
||||
expr: |
|
||||
(
|
||||
slo:http_availability:burn_rate_1h > 14.4
|
||||
and
|
||||
slo:http_availability:burn_rate_5m > 14.4
|
||||
)
|
||||
or
|
||||
(
|
||||
slo:http_availability:burn_rate_6h > 6
|
||||
and
|
||||
slo:http_availability:burn_rate_30m > 6
|
||||
)
|
||||
labels:
|
||||
severity: critical
|
||||
```
|
||||
|
||||
## SLO Review Process
|
||||
|
||||
### Weekly Review
|
||||
- Current SLO compliance
|
||||
- Error budget status
|
||||
- Trend analysis
|
||||
- Incident impact
|
||||
|
||||
### Monthly Review
|
||||
- SLO achievement
|
||||
- Error budget usage
|
||||
- Incident postmortems
|
||||
- SLO adjustments
|
||||
|
||||
### Quarterly Review
|
||||
- SLO relevance
|
||||
- Target adjustments
|
||||
- Process improvements
|
||||
- Tooling enhancements
|
||||
|
||||
## Best Practices
|
||||
|
||||
1. **Start with user-facing services**
|
||||
2. **Use multiple SLIs** (availability, latency, etc.)
|
||||
3. **Set achievable SLOs** (don't aim for 100%)
|
||||
4. **Implement multi-window alerts** to reduce noise
|
||||
5. **Track error budget** consistently
|
||||
6. **Review SLOs regularly**
|
||||
7. **Document SLO decisions**
|
||||
8. **Align with business goals**
|
||||
9. **Automate SLO reporting**
|
||||
10. **Use SLOs for prioritization**
|
||||
|
||||
## Reference Files
|
||||
|
||||
- `assets/slo-template.md` - SLO definition template
|
||||
- `references/slo-definitions.md` - SLO definition patterns
|
||||
- `references/error-budget.md` - Error budget calculations
|
||||
|
||||
## Related Skills
|
||||
|
||||
- `prometheus-configuration` - For metric collection
|
||||
- `grafana-dashboards` - For SLO visualization
|
||||
|
||||
## Limitations
|
||||
- Use this skill only when the task clearly matches the scope described above.
|
||||
- Do not treat the output as a substitute for environment-specific validation, testing, or expert review.
|
||||
- Stop and ask for clarification if required inputs, permissions, safety boundaries, or success criteria are missing.
|
||||
Reference in New Issue
Block a user