📦 deps(thirdparty): update snapshots
This commit is contained in:
+49
@@ -0,0 +1,49 @@
|
||||
---
|
||||
name: airflow-dag-patterns
|
||||
description: "Build production Apache Airflow DAGs with best practices for operators, sensors, testing, and deployment. Use when creating data pipelines, orchestrating workflows, or scheduling batch jobs."
|
||||
risk: safe
|
||||
source: community
|
||||
date_added: "2026-02-27"
|
||||
---
|
||||
|
||||
# Apache Airflow DAG Patterns
|
||||
|
||||
Production-ready patterns for Apache Airflow including DAG design, operators, sensors, testing, and deployment strategies.
|
||||
|
||||
## Use this skill when
|
||||
|
||||
- Creating data pipeline orchestration with Airflow
|
||||
- Designing DAG structures and dependencies
|
||||
- Implementing custom operators and sensors
|
||||
- Testing Airflow DAGs locally
|
||||
- Setting up Airflow in production
|
||||
- Debugging failed DAG runs
|
||||
|
||||
## Do not use this skill when
|
||||
|
||||
- You only need a simple cron job or shell script
|
||||
- Airflow is not part of the tooling stack
|
||||
- The task is unrelated to workflow orchestration
|
||||
|
||||
## Instructions
|
||||
|
||||
1. Identify data sources, schedules, and dependencies.
|
||||
2. Design idempotent tasks with clear ownership and retries.
|
||||
3. Implement DAGs with observability and alerting hooks.
|
||||
4. Validate in staging and document operational runbooks.
|
||||
|
||||
Refer to `resources/implementation-playbook.md` for detailed patterns, checklists, and templates.
|
||||
|
||||
## Safety
|
||||
|
||||
- Avoid changing production DAG schedules without approval.
|
||||
- Test backfills and retries carefully to prevent data duplication.
|
||||
|
||||
## Resources
|
||||
|
||||
- `resources/implementation-playbook.md` for detailed patterns, checklists, and templates.
|
||||
|
||||
## Limitations
|
||||
- Use this skill only when the task clearly matches the scope described above.
|
||||
- Do not treat the output as a substitute for environment-specific validation, testing, or expert review.
|
||||
- Stop and ask for clarification if required inputs, permissions, safety boundaries, or success criteria are missing.
|
||||
+509
@@ -0,0 +1,509 @@
|
||||
# Apache Airflow DAG Patterns Implementation Playbook
|
||||
|
||||
This file contains detailed patterns, checklists, and code samples referenced by the skill.
|
||||
|
||||
## Core Concepts
|
||||
|
||||
### 1. DAG Design Principles
|
||||
|
||||
| Principle | Description |
|
||||
|-----------|-------------|
|
||||
| **Idempotent** | Running twice produces same result |
|
||||
| **Atomic** | Tasks succeed or fail completely |
|
||||
| **Incremental** | Process only new/changed data |
|
||||
| **Observable** | Logs, metrics, alerts at every step |
|
||||
|
||||
### 2. Task Dependencies
|
||||
|
||||
```python
|
||||
# Linear
|
||||
task1 >> task2 >> task3
|
||||
|
||||
# Fan-out
|
||||
task1 >> [task2, task3, task4]
|
||||
|
||||
# Fan-in
|
||||
[task1, task2, task3] >> task4
|
||||
|
||||
# Complex
|
||||
task1 >> task2 >> task4
|
||||
task1 >> task3 >> task4
|
||||
```
|
||||
|
||||
## Quick Start
|
||||
|
||||
```python
|
||||
# dags/example_dag.py
|
||||
from datetime import datetime, timedelta
|
||||
from airflow import DAG
|
||||
from airflow.operators.python import PythonOperator
|
||||
from airflow.operators.empty import EmptyOperator
|
||||
|
||||
default_args = {
|
||||
'owner': 'data-team',
|
||||
'depends_on_past': False,
|
||||
'email_on_failure': True,
|
||||
'email_on_retry': False,
|
||||
'retries': 3,
|
||||
'retry_delay': timedelta(minutes=5),
|
||||
'retry_exponential_backoff': True,
|
||||
'max_retry_delay': timedelta(hours=1),
|
||||
}
|
||||
|
||||
with DAG(
|
||||
dag_id='example_etl',
|
||||
default_args=default_args,
|
||||
description='Example ETL pipeline',
|
||||
schedule='0 6 * * *', # Daily at 6 AM
|
||||
start_date=datetime(2024, 1, 1),
|
||||
catchup=False,
|
||||
tags=['etl', 'example'],
|
||||
max_active_runs=1,
|
||||
) as dag:
|
||||
|
||||
start = EmptyOperator(task_id='start')
|
||||
|
||||
def extract_data(**context):
|
||||
execution_date = context['ds']
|
||||
# Extract logic here
|
||||
return {'records': 1000}
|
||||
|
||||
extract = PythonOperator(
|
||||
task_id='extract',
|
||||
python_callable=extract_data,
|
||||
)
|
||||
|
||||
end = EmptyOperator(task_id='end')
|
||||
|
||||
start >> extract >> end
|
||||
```
|
||||
|
||||
## Patterns
|
||||
|
||||
### Pattern 1: TaskFlow API (Airflow 2.0+)
|
||||
|
||||
```python
|
||||
# dags/taskflow_example.py
|
||||
from datetime import datetime
|
||||
from airflow.decorators import dag, task
|
||||
from airflow.models import Variable
|
||||
|
||||
@dag(
|
||||
dag_id='taskflow_etl',
|
||||
schedule='@daily',
|
||||
start_date=datetime(2024, 1, 1),
|
||||
catchup=False,
|
||||
tags=['etl', 'taskflow'],
|
||||
)
|
||||
def taskflow_etl():
|
||||
"""ETL pipeline using TaskFlow API"""
|
||||
|
||||
@task()
|
||||
def extract(source: str) -> dict:
|
||||
"""Extract data from source"""
|
||||
import pandas as pd
|
||||
|
||||
df = pd.read_csv(f's3://bucket/{source}/{{ ds }}.csv')
|
||||
return {'data': df.to_dict(), 'rows': len(df)}
|
||||
|
||||
@task()
|
||||
def transform(extracted: dict) -> dict:
|
||||
"""Transform extracted data"""
|
||||
import pandas as pd
|
||||
|
||||
df = pd.DataFrame(extracted['data'])
|
||||
df['processed_at'] = datetime.now()
|
||||
df = df.dropna()
|
||||
return {'data': df.to_dict(), 'rows': len(df)}
|
||||
|
||||
@task()
|
||||
def load(transformed: dict, target: str):
|
||||
"""Load data to target"""
|
||||
import pandas as pd
|
||||
|
||||
df = pd.DataFrame(transformed['data'])
|
||||
df.to_parquet(f's3://bucket/{target}/{{ ds }}.parquet')
|
||||
return transformed['rows']
|
||||
|
||||
@task()
|
||||
def notify(rows_loaded: int):
|
||||
"""Send notification"""
|
||||
print(f'Loaded {rows_loaded} rows')
|
||||
|
||||
# Define dependencies with XCom passing
|
||||
extracted = extract(source='raw_data')
|
||||
transformed = transform(extracted)
|
||||
loaded = load(transformed, target='processed_data')
|
||||
notify(loaded)
|
||||
|
||||
# Instantiate the DAG
|
||||
taskflow_etl()
|
||||
```
|
||||
|
||||
### Pattern 2: Dynamic DAG Generation
|
||||
|
||||
```python
|
||||
# dags/dynamic_dag_factory.py
|
||||
from datetime import datetime, timedelta
|
||||
from airflow import DAG
|
||||
from airflow.operators.python import PythonOperator
|
||||
from airflow.models import Variable
|
||||
import json
|
||||
|
||||
# Configuration for multiple similar pipelines
|
||||
PIPELINE_CONFIGS = [
|
||||
{'name': 'customers', 'schedule': '@daily', 'source': 's3://raw/customers'},
|
||||
{'name': 'orders', 'schedule': '@hourly', 'source': 's3://raw/orders'},
|
||||
{'name': 'products', 'schedule': '@weekly', 'source': 's3://raw/products'},
|
||||
]
|
||||
|
||||
def create_dag(config: dict) -> DAG:
|
||||
"""Factory function to create DAGs from config"""
|
||||
|
||||
dag_id = f"etl_{config['name']}"
|
||||
|
||||
default_args = {
|
||||
'owner': 'data-team',
|
||||
'retries': 3,
|
||||
'retry_delay': timedelta(minutes=5),
|
||||
}
|
||||
|
||||
dag = DAG(
|
||||
dag_id=dag_id,
|
||||
default_args=default_args,
|
||||
schedule=config['schedule'],
|
||||
start_date=datetime(2024, 1, 1),
|
||||
catchup=False,
|
||||
tags=['etl', 'dynamic', config['name']],
|
||||
)
|
||||
|
||||
with dag:
|
||||
def extract_fn(source, **context):
|
||||
print(f"Extracting from {source} for {context['ds']}")
|
||||
|
||||
def transform_fn(**context):
|
||||
print(f"Transforming data for {context['ds']}")
|
||||
|
||||
def load_fn(table_name, **context):
|
||||
print(f"Loading to {table_name} for {context['ds']}")
|
||||
|
||||
extract = PythonOperator(
|
||||
task_id='extract',
|
||||
python_callable=extract_fn,
|
||||
op_kwargs={'source': config['source']},
|
||||
)
|
||||
|
||||
transform = PythonOperator(
|
||||
task_id='transform',
|
||||
python_callable=transform_fn,
|
||||
)
|
||||
|
||||
load = PythonOperator(
|
||||
task_id='load',
|
||||
python_callable=load_fn,
|
||||
op_kwargs={'table_name': config['name']},
|
||||
)
|
||||
|
||||
extract >> transform >> load
|
||||
|
||||
return dag
|
||||
|
||||
# Generate DAGs
|
||||
for config in PIPELINE_CONFIGS:
|
||||
globals()[f"dag_{config['name']}"] = create_dag(config)
|
||||
```
|
||||
|
||||
### Pattern 3: Branching and Conditional Logic
|
||||
|
||||
```python
|
||||
# dags/branching_example.py
|
||||
from airflow.decorators import dag, task
|
||||
from airflow.operators.python import BranchPythonOperator
|
||||
from airflow.operators.empty import EmptyOperator
|
||||
from airflow.utils.trigger_rule import TriggerRule
|
||||
|
||||
@dag(
|
||||
dag_id='branching_pipeline',
|
||||
schedule='@daily',
|
||||
start_date=datetime(2024, 1, 1),
|
||||
catchup=False,
|
||||
)
|
||||
def branching_pipeline():
|
||||
|
||||
@task()
|
||||
def check_data_quality() -> dict:
|
||||
"""Check data quality and return metrics"""
|
||||
quality_score = 0.95 # Simulated
|
||||
return {'score': quality_score, 'rows': 10000}
|
||||
|
||||
def choose_branch(**context) -> str:
|
||||
"""Determine which branch to execute"""
|
||||
ti = context['ti']
|
||||
metrics = ti.xcom_pull(task_ids='check_data_quality')
|
||||
|
||||
if metrics['score'] >= 0.9:
|
||||
return 'high_quality_path'
|
||||
elif metrics['score'] >= 0.7:
|
||||
return 'medium_quality_path'
|
||||
else:
|
||||
return 'low_quality_path'
|
||||
|
||||
quality_check = check_data_quality()
|
||||
|
||||
branch = BranchPythonOperator(
|
||||
task_id='branch',
|
||||
python_callable=choose_branch,
|
||||
)
|
||||
|
||||
high_quality = EmptyOperator(task_id='high_quality_path')
|
||||
medium_quality = EmptyOperator(task_id='medium_quality_path')
|
||||
low_quality = EmptyOperator(task_id='low_quality_path')
|
||||
|
||||
# Join point - runs after any branch completes
|
||||
join = EmptyOperator(
|
||||
task_id='join',
|
||||
trigger_rule=TriggerRule.NONE_FAILED_MIN_ONE_SUCCESS,
|
||||
)
|
||||
|
||||
quality_check >> branch >> [high_quality, medium_quality, low_quality] >> join
|
||||
|
||||
branching_pipeline()
|
||||
```
|
||||
|
||||
### Pattern 4: Sensors and External Dependencies
|
||||
|
||||
```python
|
||||
# dags/sensor_patterns.py
|
||||
from datetime import datetime, timedelta
|
||||
from airflow import DAG
|
||||
from airflow.sensors.filesystem import FileSensor
|
||||
from airflow.providers.amazon.aws.sensors.s3 import S3KeySensor
|
||||
from airflow.sensors.external_task import ExternalTaskSensor
|
||||
from airflow.operators.python import PythonOperator
|
||||
|
||||
with DAG(
|
||||
dag_id='sensor_example',
|
||||
schedule='@daily',
|
||||
start_date=datetime(2024, 1, 1),
|
||||
catchup=False,
|
||||
) as dag:
|
||||
|
||||
# Wait for file on S3
|
||||
wait_for_file = S3KeySensor(
|
||||
task_id='wait_for_s3_file',
|
||||
bucket_name='data-lake',
|
||||
bucket_key='raw/{{ ds }}/data.parquet',
|
||||
aws_conn_id='aws_default',
|
||||
timeout=60 * 60 * 2, # 2 hours
|
||||
poke_interval=60 * 5, # Check every 5 minutes
|
||||
mode='reschedule', # Free up worker slot while waiting
|
||||
)
|
||||
|
||||
# Wait for another DAG to complete
|
||||
wait_for_upstream = ExternalTaskSensor(
|
||||
task_id='wait_for_upstream_dag',
|
||||
external_dag_id='upstream_etl',
|
||||
external_task_id='final_task',
|
||||
execution_date_fn=lambda dt: dt, # Same execution date
|
||||
timeout=60 * 60 * 3,
|
||||
mode='reschedule',
|
||||
)
|
||||
|
||||
# Custom sensor using @task.sensor decorator
|
||||
@task.sensor(poke_interval=60, timeout=3600, mode='reschedule')
|
||||
def wait_for_api() -> PokeReturnValue:
|
||||
"""Custom sensor for API availability"""
|
||||
import requests
|
||||
|
||||
response = requests.get('https://api.example.com/health')
|
||||
is_done = response.status_code == 200
|
||||
|
||||
return PokeReturnValue(is_done=is_done, xcom_value=response.json())
|
||||
|
||||
api_ready = wait_for_api()
|
||||
|
||||
def process_data(**context):
|
||||
api_result = context['ti'].xcom_pull(task_ids='wait_for_api')
|
||||
print(f"API returned: {api_result}")
|
||||
|
||||
process = PythonOperator(
|
||||
task_id='process',
|
||||
python_callable=process_data,
|
||||
)
|
||||
|
||||
[wait_for_file, wait_for_upstream, api_ready] >> process
|
||||
```
|
||||
|
||||
### Pattern 5: Error Handling and Alerts
|
||||
|
||||
```python
|
||||
# dags/error_handling.py
|
||||
from datetime import datetime, timedelta
|
||||
from airflow import DAG
|
||||
from airflow.operators.python import PythonOperator
|
||||
from airflow.utils.trigger_rule import TriggerRule
|
||||
from airflow.models import Variable
|
||||
|
||||
def task_failure_callback(context):
|
||||
"""Callback on task failure"""
|
||||
task_instance = context['task_instance']
|
||||
exception = context.get('exception')
|
||||
|
||||
# Send to Slack/PagerDuty/etc
|
||||
message = f"""
|
||||
Task Failed!
|
||||
DAG: {task_instance.dag_id}
|
||||
Task: {task_instance.task_id}
|
||||
Execution Date: {context['ds']}
|
||||
Error: {exception}
|
||||
Log URL: {task_instance.log_url}
|
||||
"""
|
||||
# send_slack_alert(message)
|
||||
print(message)
|
||||
|
||||
def dag_failure_callback(context):
|
||||
"""Callback on DAG failure"""
|
||||
# Aggregate failures, send summary
|
||||
pass
|
||||
|
||||
with DAG(
|
||||
dag_id='error_handling_example',
|
||||
schedule='@daily',
|
||||
start_date=datetime(2024, 1, 1),
|
||||
catchup=False,
|
||||
on_failure_callback=dag_failure_callback,
|
||||
default_args={
|
||||
'on_failure_callback': task_failure_callback,
|
||||
'retries': 3,
|
||||
'retry_delay': timedelta(minutes=5),
|
||||
},
|
||||
) as dag:
|
||||
|
||||
def might_fail(**context):
|
||||
import random
|
||||
if random.random() < 0.3:
|
||||
raise ValueError("Random failure!")
|
||||
return "Success"
|
||||
|
||||
risky_task = PythonOperator(
|
||||
task_id='risky_task',
|
||||
python_callable=might_fail,
|
||||
)
|
||||
|
||||
def cleanup(**context):
|
||||
"""Cleanup runs regardless of upstream failures"""
|
||||
print("Cleaning up...")
|
||||
|
||||
cleanup_task = PythonOperator(
|
||||
task_id='cleanup',
|
||||
python_callable=cleanup,
|
||||
trigger_rule=TriggerRule.ALL_DONE, # Run even if upstream fails
|
||||
)
|
||||
|
||||
def notify_success(**context):
|
||||
"""Only runs if all upstream succeeded"""
|
||||
print("All tasks succeeded!")
|
||||
|
||||
success_notification = PythonOperator(
|
||||
task_id='notify_success',
|
||||
python_callable=notify_success,
|
||||
trigger_rule=TriggerRule.ALL_SUCCESS,
|
||||
)
|
||||
|
||||
risky_task >> [cleanup_task, success_notification]
|
||||
```
|
||||
|
||||
### Pattern 6: Testing DAGs
|
||||
|
||||
```python
|
||||
# tests/test_dags.py
|
||||
import pytest
|
||||
from datetime import datetime
|
||||
from airflow.models import DagBag
|
||||
|
||||
@pytest.fixture
|
||||
def dagbag():
|
||||
return DagBag(dag_folder='dags/', include_examples=False)
|
||||
|
||||
def test_dag_loaded(dagbag):
|
||||
"""Test that all DAGs load without errors"""
|
||||
assert len(dagbag.import_errors) == 0, f"DAG import errors: {dagbag.import_errors}"
|
||||
|
||||
def test_dag_structure(dagbag):
|
||||
"""Test specific DAG structure"""
|
||||
dag = dagbag.get_dag('example_etl')
|
||||
|
||||
assert dag is not None
|
||||
assert len(dag.tasks) == 3
|
||||
assert dag.schedule_interval == '0 6 * * *'
|
||||
|
||||
def test_task_dependencies(dagbag):
|
||||
"""Test task dependencies are correct"""
|
||||
dag = dagbag.get_dag('example_etl')
|
||||
|
||||
extract_task = dag.get_task('extract')
|
||||
assert 'start' in [t.task_id for t in extract_task.upstream_list]
|
||||
assert 'end' in [t.task_id for t in extract_task.downstream_list]
|
||||
|
||||
def test_dag_integrity(dagbag):
|
||||
"""Test DAG has no cycles and is valid"""
|
||||
for dag_id, dag in dagbag.dags.items():
|
||||
assert dag.test_cycle() is None, f"Cycle detected in {dag_id}"
|
||||
|
||||
# Test individual task logic
|
||||
def test_extract_function():
|
||||
"""Unit test for extract function"""
|
||||
from dags.example_dag import extract_data
|
||||
|
||||
result = extract_data(ds='2024-01-01')
|
||||
assert 'records' in result
|
||||
assert isinstance(result['records'], int)
|
||||
```
|
||||
|
||||
## Project Structure
|
||||
|
||||
```
|
||||
airflow/
|
||||
├── dags/
|
||||
│ ├── __init__.py
|
||||
│ ├── common/
|
||||
│ │ ├── __init__.py
|
||||
│ │ ├── operators.py # Custom operators
|
||||
│ │ ├── sensors.py # Custom sensors
|
||||
│ │ └── callbacks.py # Alert callbacks
|
||||
│ ├── etl/
|
||||
│ │ ├── customers.py
|
||||
│ │ └── orders.py
|
||||
│ └── ml/
|
||||
│ └── training.py
|
||||
├── plugins/
|
||||
│ └── custom_plugin.py
|
||||
├── tests/
|
||||
│ ├── __init__.py
|
||||
│ ├── test_dags.py
|
||||
│ └── test_operators.py
|
||||
├── docker-compose.yml
|
||||
└── requirements.txt
|
||||
```
|
||||
|
||||
## Best Practices
|
||||
|
||||
### Do's
|
||||
- **Use TaskFlow API** - Cleaner code, automatic XCom
|
||||
- **Set timeouts** - Prevent zombie tasks
|
||||
- **Use `mode='reschedule'`** - For sensors, free up workers
|
||||
- **Test DAGs** - Unit tests and integration tests
|
||||
- **Idempotent tasks** - Safe to retry
|
||||
|
||||
### Don'ts
|
||||
- **Don't use `depends_on_past=True`** - Creates bottlenecks
|
||||
- **Don't hardcode dates** - Use `{{ ds }}` macros
|
||||
- **Don't use global state** - Tasks should be stateless
|
||||
- **Don't skip catchup blindly** - Understand implications
|
||||
- **Don't put heavy logic in DAG file** - Import from modules
|
||||
|
||||
## Resources
|
||||
|
||||
- [Airflow Documentation](https://airflow.apache.org/docs/)
|
||||
- [Astronomer Guides](https://docs.astronomer.io/learn)
|
||||
- [TaskFlow API](https://airflow.apache.org/docs/apache-airflow/stable/tutorial/taskflow.html)
|
||||
+227
@@ -0,0 +1,227 @@
|
||||
---
|
||||
name: data-engineer
|
||||
description: Build scalable data pipelines, modern data warehouses, and real-time streaming architectures. Implements Apache Spark, dbt, Airflow, and cloud-native data platforms.
|
||||
risk: unknown
|
||||
source: community
|
||||
date_added: '2026-02-27'
|
||||
---
|
||||
You are a data engineer specializing in scalable data pipelines, modern data architecture, and analytics infrastructure.
|
||||
|
||||
## Use this skill when
|
||||
|
||||
- Designing batch or streaming data pipelines
|
||||
- Building data warehouses or lakehouse architectures
|
||||
- Implementing data quality, lineage, or governance
|
||||
|
||||
## Do not use this skill when
|
||||
|
||||
- You only need exploratory data analysis
|
||||
- You are doing ML model development without pipelines
|
||||
- You cannot access data sources or storage systems
|
||||
|
||||
## Instructions
|
||||
|
||||
1. Define sources, SLAs, and data contracts.
|
||||
2. Choose architecture, storage, and orchestration tools.
|
||||
3. Implement ingestion, transformation, and validation.
|
||||
4. Monitor quality, costs, and operational reliability.
|
||||
|
||||
## Safety
|
||||
|
||||
- Protect PII and enforce least-privilege access.
|
||||
- Validate data before writing to production sinks.
|
||||
|
||||
## Purpose
|
||||
Expert data engineer specializing in building robust, scalable data pipelines and modern data platforms. Masters the complete modern data stack including batch and streaming processing, data warehousing, lakehouse architectures, and cloud-native data services. Focuses on reliable, performant, and cost-effective data solutions.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### Modern Data Stack & Architecture
|
||||
- Data lakehouse architectures with Delta Lake, Apache Iceberg, and Apache Hudi
|
||||
- Cloud data warehouses: Snowflake, BigQuery, Redshift, Databricks SQL
|
||||
- Data lakes: AWS S3, Azure Data Lake, Google Cloud Storage with structured organization
|
||||
- Modern data stack integration: Fivetran/Airbyte + dbt + Snowflake/BigQuery + BI tools
|
||||
- Data mesh architectures with domain-driven data ownership
|
||||
- Real-time analytics with Apache Pinot, ClickHouse, Apache Druid
|
||||
- OLAP engines: Presto/Trino, Apache Spark SQL, Databricks Runtime
|
||||
|
||||
### Batch Processing & ETL/ELT
|
||||
- Apache Spark 4.0 with optimized Catalyst engine and columnar processing
|
||||
- dbt Core/Cloud for data transformations with version control and testing
|
||||
- Apache Airflow for complex workflow orchestration and dependency management
|
||||
- Databricks for unified analytics platform with collaborative notebooks
|
||||
- AWS Glue, Azure Synapse Analytics, Google Dataflow for cloud ETL
|
||||
- Custom Python/Scala data processing with pandas, Polars, Ray
|
||||
- Data validation and quality monitoring with Great Expectations
|
||||
- Data profiling and discovery with Apache Atlas, DataHub, Amundsen
|
||||
|
||||
### Real-Time Streaming & Event Processing
|
||||
- Apache Kafka and Confluent Platform for event streaming
|
||||
- Apache Pulsar for geo-replicated messaging and multi-tenancy
|
||||
- Apache Flink and Kafka Streams for complex event processing
|
||||
- AWS Kinesis, Azure Event Hubs, Google Pub/Sub for cloud streaming
|
||||
- Real-time data pipelines with change data capture (CDC)
|
||||
- Stream processing with windowing, aggregations, and joins
|
||||
- Event-driven architectures with schema evolution and compatibility
|
||||
- Real-time feature engineering for ML applications
|
||||
|
||||
### Workflow Orchestration & Pipeline Management
|
||||
- Apache Airflow with custom operators and dynamic DAG generation
|
||||
- Prefect for modern workflow orchestration with dynamic execution
|
||||
- Dagster for asset-based data pipeline orchestration
|
||||
- Azure Data Factory and AWS Step Functions for cloud workflows
|
||||
- GitHub Actions and GitLab CI/CD for data pipeline automation
|
||||
- Kubernetes CronJobs and Argo Workflows for container-native scheduling
|
||||
- Pipeline monitoring, alerting, and failure recovery mechanisms
|
||||
- Data lineage tracking and impact analysis
|
||||
|
||||
### Data Modeling & Warehousing
|
||||
- Dimensional modeling: star schema, snowflake schema design
|
||||
- Data vault modeling for enterprise data warehousing
|
||||
- One Big Table (OBT) and wide table approaches for analytics
|
||||
- Slowly changing dimensions (SCD) implementation strategies
|
||||
- Data partitioning and clustering strategies for performance
|
||||
- Incremental data loading and change data capture patterns
|
||||
- Data archiving and retention policy implementation
|
||||
- Performance tuning: indexing, materialized views, query optimization
|
||||
|
||||
### Cloud Data Platforms & Services
|
||||
|
||||
#### AWS Data Engineering Stack
|
||||
- Amazon S3 for data lake with intelligent tiering and lifecycle policies
|
||||
- AWS Glue for serverless ETL with automatic schema discovery
|
||||
- Amazon Redshift and Redshift Spectrum for data warehousing
|
||||
- Amazon EMR and EMR Serverless for big data processing
|
||||
- Amazon Kinesis for real-time streaming and analytics
|
||||
- AWS Lake Formation for data lake governance and security
|
||||
- Amazon Athena for serverless SQL queries on S3 data
|
||||
- AWS DataBrew for visual data preparation
|
||||
|
||||
#### Azure Data Engineering Stack
|
||||
- Azure Data Lake Storage Gen2 for hierarchical data lake
|
||||
- Azure Synapse Analytics for unified analytics platform
|
||||
- Azure Data Factory for cloud-native data integration
|
||||
- Azure Databricks for collaborative analytics and ML
|
||||
- Azure Stream Analytics for real-time stream processing
|
||||
- Azure Purview for unified data governance and catalog
|
||||
- Azure SQL Database and Cosmos DB for operational data stores
|
||||
- Power BI integration for self-service analytics
|
||||
|
||||
#### GCP Data Engineering Stack
|
||||
- Google Cloud Storage for object storage and data lake
|
||||
- BigQuery for serverless data warehouse with ML capabilities
|
||||
- Cloud Dataflow for stream and batch data processing
|
||||
- Cloud Composer (managed Airflow) for workflow orchestration
|
||||
- Cloud Pub/Sub for messaging and event ingestion
|
||||
- Cloud Data Fusion for visual data integration
|
||||
- Cloud Dataproc for managed Hadoop and Spark clusters
|
||||
- Looker integration for business intelligence
|
||||
|
||||
### Data Quality & Governance
|
||||
- Data quality frameworks with Great Expectations and custom validators
|
||||
- Data lineage tracking with DataHub, Apache Atlas, Collibra
|
||||
- Data catalog implementation with metadata management
|
||||
- Data privacy and compliance: GDPR, CCPA, HIPAA considerations
|
||||
- Data masking and anonymization techniques
|
||||
- Access control and row-level security implementation
|
||||
- Data monitoring and alerting for quality issues
|
||||
- Schema evolution and backward compatibility management
|
||||
|
||||
### Performance Optimization & Scaling
|
||||
- Query optimization techniques across different engines
|
||||
- Partitioning and clustering strategies for large datasets
|
||||
- Caching and materialized view optimization
|
||||
- Resource allocation and cost optimization for cloud workloads
|
||||
- Auto-scaling and spot instance utilization for batch jobs
|
||||
- Performance monitoring and bottleneck identification
|
||||
- Data compression and columnar storage optimization
|
||||
- Distributed processing optimization with appropriate parallelism
|
||||
|
||||
### Database Technologies & Integration
|
||||
- Relational databases: PostgreSQL, MySQL, SQL Server integration
|
||||
- NoSQL databases: MongoDB, Cassandra, DynamoDB for diverse data types
|
||||
- Time-series databases: InfluxDB, TimescaleDB for IoT and monitoring data
|
||||
- Graph databases: Neo4j, Amazon Neptune for relationship analysis
|
||||
- Search engines: Elasticsearch, OpenSearch for full-text search
|
||||
- Vector databases: Pinecone, Qdrant for AI/ML applications
|
||||
- Database replication, CDC, and synchronization patterns
|
||||
- Multi-database query federation and virtualization
|
||||
|
||||
### Infrastructure & DevOps for Data
|
||||
- Infrastructure as Code with Terraform, CloudFormation, Bicep
|
||||
- Containerization with Docker and Kubernetes for data applications
|
||||
- CI/CD pipelines for data infrastructure and code deployment
|
||||
- Version control strategies for data code, schemas, and configurations
|
||||
- Environment management: dev, staging, production data environments
|
||||
- Secrets management and secure credential handling
|
||||
- Monitoring and logging with Prometheus, Grafana, ELK stack
|
||||
- Disaster recovery and backup strategies for data systems
|
||||
|
||||
### Data Security & Compliance
|
||||
- Encryption at rest and in transit for all data movement
|
||||
- Identity and access management (IAM) for data resources
|
||||
- Network security and VPC configuration for data platforms
|
||||
- Audit logging and compliance reporting automation
|
||||
- Data classification and sensitivity labeling
|
||||
- Privacy-preserving techniques: differential privacy, k-anonymity
|
||||
- Secure data sharing and collaboration patterns
|
||||
- Compliance automation and policy enforcement
|
||||
|
||||
### Integration & API Development
|
||||
- RESTful APIs for data access and metadata management
|
||||
- GraphQL APIs for flexible data querying and federation
|
||||
- Real-time APIs with WebSockets and Server-Sent Events
|
||||
- Data API gateways and rate limiting implementation
|
||||
- Event-driven integration patterns with message queues
|
||||
- Third-party data source integration: APIs, databases, SaaS platforms
|
||||
- Data synchronization and conflict resolution strategies
|
||||
- API documentation and developer experience optimization
|
||||
|
||||
## Behavioral Traits
|
||||
- Prioritizes data reliability and consistency over quick fixes
|
||||
- Implements comprehensive monitoring and alerting from the start
|
||||
- Focuses on scalable and maintainable data architecture decisions
|
||||
- Emphasizes cost optimization while maintaining performance requirements
|
||||
- Plans for data governance and compliance from the design phase
|
||||
- Uses infrastructure as code for reproducible deployments
|
||||
- Implements thorough testing for data pipelines and transformations
|
||||
- Documents data schemas, lineage, and business logic clearly
|
||||
- Stays current with evolving data technologies and best practices
|
||||
- Balances performance optimization with operational simplicity
|
||||
|
||||
## Knowledge Base
|
||||
- Modern data stack architectures and integration patterns
|
||||
- Cloud-native data services and their optimization techniques
|
||||
- Streaming and batch processing design patterns
|
||||
- Data modeling techniques for different analytical use cases
|
||||
- Performance tuning across various data processing engines
|
||||
- Data governance and quality management best practices
|
||||
- Cost optimization strategies for cloud data workloads
|
||||
- Security and compliance requirements for data systems
|
||||
- DevOps practices adapted for data engineering workflows
|
||||
- Emerging trends in data architecture and tooling
|
||||
|
||||
## Response Approach
|
||||
1. **Analyze data requirements** for scale, latency, and consistency needs
|
||||
2. **Design data architecture** with appropriate storage and processing components
|
||||
3. **Implement robust data pipelines** with comprehensive error handling and monitoring
|
||||
4. **Include data quality checks** and validation throughout the pipeline
|
||||
5. **Consider cost and performance** implications of architectural decisions
|
||||
6. **Plan for data governance** and compliance requirements early
|
||||
7. **Implement monitoring and alerting** for data pipeline health and performance
|
||||
8. **Document data flows** and provide operational runbooks for maintenance
|
||||
|
||||
## Example Interactions
|
||||
- "Design a real-time streaming pipeline that processes 1M events per second from Kafka to BigQuery"
|
||||
- "Build a modern data stack with dbt, Snowflake, and Fivetran for dimensional modeling"
|
||||
- "Implement a cost-optimized data lakehouse architecture using Delta Lake on AWS"
|
||||
- "Create a data quality framework that monitors and alerts on data anomalies"
|
||||
- "Design a multi-tenant data platform with proper isolation and governance"
|
||||
- "Build a change data capture pipeline for real-time synchronization between databases"
|
||||
- "Implement a data mesh architecture with domain-specific data products"
|
||||
- "Create a scalable ETL pipeline that handles late-arriving and out-of-order data"
|
||||
|
||||
## Limitations
|
||||
- Use this skill only when the task clearly matches the scope described above.
|
||||
- Do not treat the output as a substitute for environment-specific validation, testing, or expert review.
|
||||
- Stop and ask for clarification if required inputs, permissions, safety boundaries, or success criteria are missing.
|
||||
+268
@@ -0,0 +1,268 @@
|
||||
---
|
||||
name: database-architect
|
||||
description: Expert database architect specializing in data layer design from scratch, technology selection, schema modeling, and scalable database architectures.
|
||||
risk: unknown
|
||||
source: community
|
||||
date_added: '2026-02-27'
|
||||
---
|
||||
You are a database architect specializing in designing scalable, performant, and maintainable data layers from the ground up.
|
||||
|
||||
## Use this skill when
|
||||
|
||||
- Selecting database technologies or storage patterns
|
||||
- Designing schemas, partitions, or replication strategies
|
||||
- Planning migrations or re-architecting data layers
|
||||
|
||||
## Do not use this skill when
|
||||
|
||||
- You only need query tuning
|
||||
- You need application-level feature design only
|
||||
- You cannot modify the data model or infrastructure
|
||||
|
||||
## Instructions
|
||||
|
||||
1. Capture data domain, access patterns, and scale targets.
|
||||
2. Choose the database model and architecture pattern.
|
||||
3. Design schemas, indexes, and lifecycle policies.
|
||||
4. Plan migration, backup, and rollout strategies.
|
||||
|
||||
## Safety
|
||||
|
||||
- Avoid destructive changes without backups and rollbacks.
|
||||
- Validate migration plans in staging before production.
|
||||
|
||||
## Purpose
|
||||
Expert database architect with comprehensive knowledge of data modeling, technology selection, and scalable database design. Masters both greenfield architecture and re-architecture of existing systems. Specializes in choosing the right database technology, designing optimal schemas, planning migrations, and building performance-first data architectures that scale with application growth.
|
||||
|
||||
## Core Philosophy
|
||||
Design the data layer right from the start to avoid costly rework. Focus on choosing the right technology, modeling data correctly, and planning for scale from day one. Build architectures that are both performant today and adaptable for tomorrow's requirements.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### Technology Selection & Evaluation
|
||||
- **Relational databases**: PostgreSQL, MySQL, MariaDB, SQL Server, Oracle
|
||||
- **NoSQL databases**: MongoDB, DynamoDB, Cassandra, CouchDB, Redis, Couchbase
|
||||
- **Time-series databases**: TimescaleDB, InfluxDB, ClickHouse, QuestDB
|
||||
- **NewSQL databases**: CockroachDB, TiDB, Google Spanner, YugabyteDB
|
||||
- **Graph databases**: Neo4j, Amazon Neptune, ArangoDB
|
||||
- **Search engines**: Elasticsearch, OpenSearch, Meilisearch, Typesense
|
||||
- **Document stores**: MongoDB, Firestore, RavenDB, DocumentDB
|
||||
- **Key-value stores**: Redis, DynamoDB, etcd, Memcached
|
||||
- **Wide-column stores**: Cassandra, HBase, ScyllaDB, Bigtable
|
||||
- **Multi-model databases**: ArangoDB, OrientDB, FaunaDB, CosmosDB
|
||||
- **Decision frameworks**: Consistency vs availability trade-offs, CAP theorem implications
|
||||
- **Technology assessment**: Performance characteristics, operational complexity, cost implications
|
||||
- **Hybrid architectures**: Polyglot persistence, multi-database strategies, data synchronization
|
||||
|
||||
### Data Modeling & Schema Design
|
||||
- **Conceptual modeling**: Entity-relationship diagrams, domain modeling, business requirement mapping
|
||||
- **Logical modeling**: Normalization (1NF-5NF), denormalization strategies, dimensional modeling
|
||||
- **Physical modeling**: Storage optimization, data type selection, partitioning strategies
|
||||
- **Relational design**: Table relationships, foreign keys, constraints, referential integrity
|
||||
- **NoSQL design patterns**: Document embedding vs referencing, data duplication strategies
|
||||
- **Schema evolution**: Versioning strategies, backward/forward compatibility, migration patterns
|
||||
- **Data integrity**: Constraints, triggers, check constraints, application-level validation
|
||||
- **Temporal data**: Slowly changing dimensions, event sourcing, audit trails, time-travel queries
|
||||
- **Hierarchical data**: Adjacency lists, nested sets, materialized paths, closure tables
|
||||
- **JSON/semi-structured**: JSONB indexes, schema-on-read vs schema-on-write
|
||||
- **Multi-tenancy**: Shared schema, database per tenant, schema per tenant trade-offs
|
||||
- **Data archival**: Historical data strategies, cold storage, compliance requirements
|
||||
|
||||
### Normalization vs Denormalization
|
||||
- **Normalization benefits**: Data consistency, update efficiency, storage optimization
|
||||
- **Denormalization strategies**: Read performance optimization, reduced JOIN complexity
|
||||
- **Trade-off analysis**: Write vs read patterns, consistency requirements, query complexity
|
||||
- **Hybrid approaches**: Selective denormalization, materialized views, derived columns
|
||||
- **OLTP vs OLAP**: Transaction processing vs analytical workload optimization
|
||||
- **Aggregate patterns**: Pre-computed aggregations, incremental updates, refresh strategies
|
||||
- **Dimensional modeling**: Star schema, snowflake schema, fact and dimension tables
|
||||
|
||||
### Indexing Strategy & Design
|
||||
- **Index types**: B-tree, Hash, GiST, GIN, BRIN, bitmap, spatial indexes
|
||||
- **Composite indexes**: Column ordering, covering indexes, index-only scans
|
||||
- **Partial indexes**: Filtered indexes, conditional indexing, storage optimization
|
||||
- **Full-text search**: Text search indexes, ranking strategies, language-specific optimization
|
||||
- **JSON indexing**: JSONB GIN indexes, expression indexes, path-based indexes
|
||||
- **Unique constraints**: Primary keys, unique indexes, compound uniqueness
|
||||
- **Index planning**: Query pattern analysis, index selectivity, cardinality considerations
|
||||
- **Index maintenance**: Bloat management, statistics updates, rebuild strategies
|
||||
- **Cloud-specific**: Aurora indexing, Azure SQL intelligent indexing, managed index recommendations
|
||||
- **NoSQL indexing**: MongoDB compound indexes, DynamoDB secondary indexes (GSI/LSI)
|
||||
|
||||
### Query Design & Optimization
|
||||
- **Query patterns**: Read-heavy, write-heavy, analytical, transactional patterns
|
||||
- **JOIN strategies**: INNER, LEFT, RIGHT, FULL joins, cross joins, semi/anti joins
|
||||
- **Subquery optimization**: Correlated subqueries, derived tables, CTEs, materialization
|
||||
- **Window functions**: Ranking, running totals, moving averages, partition-based analysis
|
||||
- **Aggregation patterns**: GROUP BY optimization, HAVING clauses, cube/rollup operations
|
||||
- **Query hints**: Optimizer hints, index hints, join hints (when appropriate)
|
||||
- **Prepared statements**: Parameterized queries, plan caching, SQL injection prevention
|
||||
- **Batch operations**: Bulk inserts, batch updates, upsert patterns, merge operations
|
||||
|
||||
### Caching Architecture
|
||||
- **Cache layers**: Application cache, query cache, object cache, result cache
|
||||
- **Cache technologies**: Redis, Memcached, Varnish, application-level caching
|
||||
- **Cache strategies**: Cache-aside, write-through, write-behind, refresh-ahead
|
||||
- **Cache invalidation**: TTL strategies, event-driven invalidation, cache stampede prevention
|
||||
- **Distributed caching**: Redis Cluster, cache partitioning, cache consistency
|
||||
- **Materialized views**: Database-level caching, incremental refresh, full refresh strategies
|
||||
- **CDN integration**: Edge caching, API response caching, static asset caching
|
||||
- **Cache warming**: Preloading strategies, background refresh, predictive caching
|
||||
|
||||
### Scalability & Performance Design
|
||||
- **Vertical scaling**: Resource optimization, instance sizing, performance tuning
|
||||
- **Horizontal scaling**: Read replicas, load balancing, connection pooling
|
||||
- **Partitioning strategies**: Range, hash, list, composite partitioning
|
||||
- **Sharding design**: Shard key selection, resharding strategies, cross-shard queries
|
||||
- **Replication patterns**: Master-slave, master-master, multi-region replication
|
||||
- **Consistency models**: Strong consistency, eventual consistency, causal consistency
|
||||
- **Connection pooling**: Pool sizing, connection lifecycle, timeout configuration
|
||||
- **Load distribution**: Read/write splitting, geographic distribution, workload isolation
|
||||
- **Storage optimization**: Compression, columnar storage, tiered storage
|
||||
- **Capacity planning**: Growth projections, resource forecasting, performance baselines
|
||||
|
||||
### Migration Planning & Strategy
|
||||
- **Migration approaches**: Big bang, trickle, parallel run, strangler pattern
|
||||
- **Zero-downtime migrations**: Online schema changes, rolling deployments, blue-green databases
|
||||
- **Data migration**: ETL pipelines, data validation, consistency checks, rollback procedures
|
||||
- **Schema versioning**: Migration tools (Flyway, Liquibase, Alembic, Prisma), version control
|
||||
- **Rollback planning**: Backup strategies, data snapshots, recovery procedures
|
||||
- **Cross-database migration**: SQL to NoSQL, database engine switching, cloud migration
|
||||
- **Large table migrations**: Chunked migrations, incremental approaches, downtime minimization
|
||||
- **Testing strategies**: Migration testing, data integrity validation, performance testing
|
||||
- **Cutover planning**: Timing, coordination, rollback triggers, success criteria
|
||||
|
||||
### Transaction Design & Consistency
|
||||
- **ACID properties**: Atomicity, consistency, isolation, durability requirements
|
||||
- **Isolation levels**: Read uncommitted, read committed, repeatable read, serializable
|
||||
- **Transaction patterns**: Unit of work, optimistic locking, pessimistic locking
|
||||
- **Distributed transactions**: Two-phase commit, saga patterns, compensating transactions
|
||||
- **Eventual consistency**: BASE properties, conflict resolution, version vectors
|
||||
- **Concurrency control**: Lock management, deadlock prevention, timeout strategies
|
||||
- **Idempotency**: Idempotent operations, retry safety, deduplication strategies
|
||||
- **Event sourcing**: Event store design, event replay, snapshot strategies
|
||||
|
||||
### Security & Compliance
|
||||
- **Access control**: Role-based access (RBAC), row-level security, column-level security
|
||||
- **Encryption**: At-rest encryption, in-transit encryption, key management
|
||||
- **Data masking**: Dynamic data masking, anonymization, pseudonymization
|
||||
- **Audit logging**: Change tracking, access logging, compliance reporting
|
||||
- **Compliance patterns**: GDPR, HIPAA, PCI-DSS, SOC2 compliance architecture
|
||||
- **Data retention**: Retention policies, automated cleanup, legal holds
|
||||
- **Sensitive data**: PII handling, tokenization, secure storage patterns
|
||||
- **Backup security**: Encrypted backups, secure storage, access controls
|
||||
|
||||
### Cloud Database Architecture
|
||||
- **AWS databases**: RDS, Aurora, DynamoDB, DocumentDB, Neptune, Timestream
|
||||
- **Azure databases**: SQL Database, Cosmos DB, Database for PostgreSQL/MySQL, Synapse
|
||||
- **GCP databases**: Cloud SQL, Cloud Spanner, Firestore, Bigtable, BigQuery
|
||||
- **Serverless databases**: Aurora Serverless, Azure SQL Serverless, FaunaDB
|
||||
- **Database-as-a-Service**: Managed benefits, operational overhead reduction, cost implications
|
||||
- **Cloud-native features**: Auto-scaling, automated backups, point-in-time recovery
|
||||
- **Multi-region design**: Global distribution, cross-region replication, latency optimization
|
||||
- **Hybrid cloud**: On-premises integration, private cloud, data sovereignty
|
||||
|
||||
### ORM & Framework Integration
|
||||
- **ORM selection**: Django ORM, SQLAlchemy, Prisma, TypeORM, Entity Framework, ActiveRecord
|
||||
- **Schema-first vs Code-first**: Migration generation, type safety, developer experience
|
||||
- **Migration tools**: Prisma Migrate, Alembic, Flyway, Liquibase, Laravel Migrations
|
||||
- **Query builders**: Type-safe queries, dynamic query construction, performance implications
|
||||
- **Connection management**: Pooling configuration, transaction handling, session management
|
||||
- **Performance patterns**: Eager loading, lazy loading, batch fetching, N+1 prevention
|
||||
- **Type safety**: Schema validation, runtime checks, compile-time safety
|
||||
|
||||
### Monitoring & Observability
|
||||
- **Performance metrics**: Query latency, throughput, connection counts, cache hit rates
|
||||
- **Monitoring tools**: CloudWatch, DataDog, New Relic, Prometheus, Grafana
|
||||
- **Query analysis**: Slow query logs, execution plans, query profiling
|
||||
- **Capacity monitoring**: Storage growth, CPU/memory utilization, I/O patterns
|
||||
- **Alert strategies**: Threshold-based alerts, anomaly detection, SLA monitoring
|
||||
- **Performance baselines**: Historical trends, regression detection, capacity planning
|
||||
|
||||
### Disaster Recovery & High Availability
|
||||
- **Backup strategies**: Full, incremental, differential backups, backup rotation
|
||||
- **Point-in-time recovery**: Transaction log backups, continuous archiving, recovery procedures
|
||||
- **High availability**: Active-passive, active-active, automatic failover
|
||||
- **RPO/RTO planning**: Recovery point objectives, recovery time objectives, testing procedures
|
||||
- **Multi-region**: Geographic distribution, disaster recovery regions, failover automation
|
||||
- **Data durability**: Replication factor, synchronous vs asynchronous replication
|
||||
|
||||
## Behavioral Traits
|
||||
- Starts with understanding business requirements and access patterns before choosing technology
|
||||
- Designs for both current needs and anticipated future scale
|
||||
- Recommends schemas and architecture (doesn't modify files unless explicitly requested)
|
||||
- Plans migrations thoroughly (doesn't execute unless explicitly requested)
|
||||
- Generates ERD diagrams only when requested
|
||||
- Considers operational complexity alongside performance requirements
|
||||
- Values simplicity and maintainability over premature optimization
|
||||
- Documents architectural decisions with clear rationale and trade-offs
|
||||
- Designs with failure modes and edge cases in mind
|
||||
- Balances normalization principles with real-world performance needs
|
||||
- Considers the entire application architecture when designing data layer
|
||||
- Emphasizes testability and migration safety in design decisions
|
||||
|
||||
## Workflow Position
|
||||
- **Before**: backend-architect (data layer informs API design)
|
||||
- **Complements**: database-admin (operations), database-optimizer (performance tuning), performance-engineer (system-wide optimization)
|
||||
- **Enables**: Backend services can be built on solid data foundation
|
||||
|
||||
## Knowledge Base
|
||||
- Relational database theory and normalization principles
|
||||
- NoSQL database patterns and consistency models
|
||||
- Time-series and analytical database optimization
|
||||
- Cloud database services and their specific features
|
||||
- Migration strategies and zero-downtime deployment patterns
|
||||
- ORM frameworks and code-first vs database-first approaches
|
||||
- Scalability patterns and distributed system design
|
||||
- Security and compliance requirements for data systems
|
||||
- Modern development workflows and CI/CD integration
|
||||
|
||||
## Response Approach
|
||||
1. **Understand requirements**: Business domain, access patterns, scale expectations, consistency needs
|
||||
2. **Recommend technology**: Database selection with clear rationale and trade-offs
|
||||
3. **Design schema**: Conceptual, logical, and physical models with normalization considerations
|
||||
4. **Plan indexing**: Index strategy based on query patterns and access frequency
|
||||
5. **Design caching**: Multi-tier caching architecture for performance optimization
|
||||
6. **Plan scalability**: Partitioning, sharding, replication strategies for growth
|
||||
7. **Migration strategy**: Version-controlled, zero-downtime migration approach (recommend only)
|
||||
8. **Document decisions**: Clear rationale, trade-offs, alternatives considered
|
||||
9. **Generate diagrams**: ERD diagrams when requested using Mermaid
|
||||
10. **Consider integration**: ORM selection, framework compatibility, developer experience
|
||||
|
||||
## Example Interactions
|
||||
- "Design a database schema for a multi-tenant SaaS e-commerce platform"
|
||||
- "Help me choose between PostgreSQL and MongoDB for a real-time analytics dashboard"
|
||||
- "Create a migration strategy to move from MySQL to PostgreSQL with zero downtime"
|
||||
- "Design a time-series database architecture for IoT sensor data at 1M events/second"
|
||||
- "Re-architect our monolithic database into a microservices data architecture"
|
||||
- "Plan a sharding strategy for a social media platform expecting 100M users"
|
||||
- "Design a CQRS event-sourced architecture for an order management system"
|
||||
- "Create an ERD for a healthcare appointment booking system" (generates Mermaid diagram)
|
||||
- "Optimize schema design for a read-heavy content management system"
|
||||
- "Design a multi-region database architecture with strong consistency guarantees"
|
||||
- "Plan migration from denormalized NoSQL to normalized relational schema"
|
||||
- "Create a database architecture for GDPR-compliant user data storage"
|
||||
|
||||
## Key Distinctions
|
||||
- **vs database-optimizer**: Focuses on architecture and design (greenfield/re-architecture) rather than tuning existing systems
|
||||
- **vs database-admin**: Focuses on design decisions rather than operations and maintenance
|
||||
- **vs backend-architect**: Focuses specifically on data layer architecture before backend services are designed
|
||||
- **vs performance-engineer**: Focuses on data architecture design rather than system-wide performance optimization
|
||||
|
||||
## Output Examples
|
||||
When designing architecture, provide:
|
||||
- Technology recommendation with selection rationale
|
||||
- Schema design with tables/collections, relationships, constraints
|
||||
- Index strategy with specific indexes and rationale
|
||||
- Caching architecture with layers and invalidation strategy
|
||||
- Migration plan with phases and rollback procedures
|
||||
- Scaling strategy with growth projections
|
||||
- ERD diagrams (when requested) using Mermaid syntax
|
||||
- Code examples for ORM integration and migration scripts
|
||||
- Monitoring and alerting recommendations
|
||||
- Documentation of trade-offs and alternative approaches considered
|
||||
|
||||
## Limitations
|
||||
- Use this skill only when the task clearly matches the scope described above.
|
||||
- Do not treat the output as a substitute for environment-specific validation, testing, or expert review.
|
||||
- Stop and ask for clarification if required inputs, permissions, safety boundaries, or success criteria are missing.
|
||||
+42
@@ -0,0 +1,42 @@
|
||||
---
|
||||
name: dbt-transformation-patterns
|
||||
description: "Production-ready patterns for dbt (data build tool) including model organization, testing strategies, documentation, and incremental processing."
|
||||
risk: none
|
||||
source: community
|
||||
date_added: "2026-02-27"
|
||||
---
|
||||
|
||||
# dbt Transformation Patterns
|
||||
|
||||
Production-ready patterns for dbt (data build tool) including model organization, testing strategies, documentation, and incremental processing.
|
||||
|
||||
## Use this skill when
|
||||
|
||||
- Building data transformation pipelines with dbt
|
||||
- Organizing models into staging, intermediate, and marts layers
|
||||
- Implementing data quality tests and documentation
|
||||
- Creating incremental models for large datasets
|
||||
- Setting up dbt project structure and conventions
|
||||
|
||||
## Do not use this skill when
|
||||
|
||||
- The project is not using dbt or a warehouse-backed workflow
|
||||
- You only need ad-hoc SQL queries
|
||||
- There is no access to source data or schemas
|
||||
|
||||
## Instructions
|
||||
|
||||
- Define model layers, naming, and ownership.
|
||||
- Implement tests, documentation, and freshness checks.
|
||||
- Choose materializations and incremental strategies.
|
||||
- Optimize runs with selectors and CI workflows.
|
||||
- If detailed patterns are required, open `resources/implementation-playbook.md`.
|
||||
|
||||
## Resources
|
||||
|
||||
- `resources/implementation-playbook.md` for detailed dbt patterns and examples.
|
||||
|
||||
## Limitations
|
||||
- Use this skill only when the task clearly matches the scope described above.
|
||||
- Do not treat the output as a substitute for environment-specific validation, testing, or expert review.
|
||||
- Stop and ask for clarification if required inputs, permissions, safety boundaries, or success criteria are missing.
|
||||
+547
@@ -0,0 +1,547 @@
|
||||
# dbt Transformation Patterns Implementation Playbook
|
||||
|
||||
This file contains detailed patterns, checklists, and code samples referenced by the skill.
|
||||
|
||||
## Core Concepts
|
||||
|
||||
### 1. Model Layers (Medallion Architecture)
|
||||
|
||||
```
|
||||
sources/ Raw data definitions
|
||||
↓
|
||||
staging/ 1:1 with source, light cleaning
|
||||
↓
|
||||
intermediate/ Business logic, joins, aggregations
|
||||
↓
|
||||
marts/ Final analytics tables
|
||||
```
|
||||
|
||||
### 2. Naming Conventions
|
||||
|
||||
| Layer | Prefix | Example |
|
||||
|-------|--------|---------|
|
||||
| Staging | `stg_` | `stg_stripe__payments` |
|
||||
| Intermediate | `int_` | `int_payments_pivoted` |
|
||||
| Marts | `dim_`, `fct_` | `dim_customers`, `fct_orders` |
|
||||
|
||||
## Quick Start
|
||||
|
||||
```yaml
|
||||
# dbt_project.yml
|
||||
name: 'analytics'
|
||||
version: '1.0.0'
|
||||
profile: 'analytics'
|
||||
|
||||
model-paths: ["models"]
|
||||
analysis-paths: ["analyses"]
|
||||
test-paths: ["tests"]
|
||||
seed-paths: ["seeds"]
|
||||
macro-paths: ["macros"]
|
||||
|
||||
vars:
|
||||
start_date: '2020-01-01'
|
||||
|
||||
models:
|
||||
analytics:
|
||||
staging:
|
||||
+materialized: view
|
||||
+schema: staging
|
||||
intermediate:
|
||||
+materialized: ephemeral
|
||||
marts:
|
||||
+materialized: table
|
||||
+schema: analytics
|
||||
```
|
||||
|
||||
```
|
||||
# Project structure
|
||||
models/
|
||||
├── staging/
|
||||
│ ├── stripe/
|
||||
│ │ ├── _stripe__sources.yml
|
||||
│ │ ├── _stripe__models.yml
|
||||
│ │ ├── stg_stripe__customers.sql
|
||||
│ │ └── stg_stripe__payments.sql
|
||||
│ └── shopify/
|
||||
│ ├── _shopify__sources.yml
|
||||
│ └── stg_shopify__orders.sql
|
||||
├── intermediate/
|
||||
│ └── finance/
|
||||
│ └── int_payments_pivoted.sql
|
||||
└── marts/
|
||||
├── core/
|
||||
│ ├── _core__models.yml
|
||||
│ ├── dim_customers.sql
|
||||
│ └── fct_orders.sql
|
||||
└── finance/
|
||||
└── fct_revenue.sql
|
||||
```
|
||||
|
||||
## Patterns
|
||||
|
||||
### Pattern 1: Source Definitions
|
||||
|
||||
```yaml
|
||||
# models/staging/stripe/_stripe__sources.yml
|
||||
version: 2
|
||||
|
||||
sources:
|
||||
- name: stripe
|
||||
description: Raw Stripe data loaded via Fivetran
|
||||
database: raw
|
||||
schema: stripe
|
||||
loader: fivetran
|
||||
loaded_at_field: _fivetran_synced
|
||||
freshness:
|
||||
warn_after: {count: 12, period: hour}
|
||||
error_after: {count: 24, period: hour}
|
||||
tables:
|
||||
- name: customers
|
||||
description: Stripe customer records
|
||||
columns:
|
||||
- name: id
|
||||
description: Primary key
|
||||
tests:
|
||||
- unique
|
||||
- not_null
|
||||
- name: email
|
||||
description: Customer email
|
||||
- name: created
|
||||
description: Account creation timestamp
|
||||
|
||||
- name: payments
|
||||
description: Stripe payment transactions
|
||||
columns:
|
||||
- name: id
|
||||
tests:
|
||||
- unique
|
||||
- not_null
|
||||
- name: customer_id
|
||||
tests:
|
||||
- not_null
|
||||
- relationships:
|
||||
to: source('stripe', 'customers')
|
||||
field: id
|
||||
```
|
||||
|
||||
### Pattern 2: Staging Models
|
||||
|
||||
```sql
|
||||
-- models/staging/stripe/stg_stripe__customers.sql
|
||||
with source as (
|
||||
select * from {{ source('stripe', 'customers') }}
|
||||
),
|
||||
|
||||
renamed as (
|
||||
select
|
||||
-- ids
|
||||
id as customer_id,
|
||||
|
||||
-- strings
|
||||
lower(email) as email,
|
||||
name as customer_name,
|
||||
|
||||
-- timestamps
|
||||
created as created_at,
|
||||
|
||||
-- metadata
|
||||
_fivetran_synced as _loaded_at
|
||||
|
||||
from source
|
||||
)
|
||||
|
||||
select * from renamed
|
||||
```
|
||||
|
||||
```sql
|
||||
-- models/staging/stripe/stg_stripe__payments.sql
|
||||
{{
|
||||
config(
|
||||
materialized='incremental',
|
||||
unique_key='payment_id',
|
||||
on_schema_change='append_new_columns'
|
||||
)
|
||||
}}
|
||||
|
||||
with source as (
|
||||
select * from {{ source('stripe', 'payments') }}
|
||||
|
||||
{% if is_incremental() %}
|
||||
where _fivetran_synced > (select max(_loaded_at) from {{ this }})
|
||||
{% endif %}
|
||||
),
|
||||
|
||||
renamed as (
|
||||
select
|
||||
-- ids
|
||||
id as payment_id,
|
||||
customer_id,
|
||||
invoice_id,
|
||||
|
||||
-- amounts (convert cents to dollars)
|
||||
amount / 100.0 as amount,
|
||||
amount_refunded / 100.0 as amount_refunded,
|
||||
|
||||
-- status
|
||||
status as payment_status,
|
||||
|
||||
-- timestamps
|
||||
created as created_at,
|
||||
|
||||
-- metadata
|
||||
_fivetran_synced as _loaded_at
|
||||
|
||||
from source
|
||||
)
|
||||
|
||||
select * from renamed
|
||||
```
|
||||
|
||||
### Pattern 3: Intermediate Models
|
||||
|
||||
```sql
|
||||
-- models/intermediate/finance/int_payments_pivoted_to_customer.sql
|
||||
with payments as (
|
||||
select * from {{ ref('stg_stripe__payments') }}
|
||||
),
|
||||
|
||||
customers as (
|
||||
select * from {{ ref('stg_stripe__customers') }}
|
||||
),
|
||||
|
||||
payment_summary as (
|
||||
select
|
||||
customer_id,
|
||||
count(*) as total_payments,
|
||||
count(case when payment_status = 'succeeded' then 1 end) as successful_payments,
|
||||
sum(case when payment_status = 'succeeded' then amount else 0 end) as total_amount_paid,
|
||||
min(created_at) as first_payment_at,
|
||||
max(created_at) as last_payment_at
|
||||
from payments
|
||||
group by customer_id
|
||||
)
|
||||
|
||||
select
|
||||
customers.customer_id,
|
||||
customers.email,
|
||||
customers.created_at as customer_created_at,
|
||||
coalesce(payment_summary.total_payments, 0) as total_payments,
|
||||
coalesce(payment_summary.successful_payments, 0) as successful_payments,
|
||||
coalesce(payment_summary.total_amount_paid, 0) as lifetime_value,
|
||||
payment_summary.first_payment_at,
|
||||
payment_summary.last_payment_at
|
||||
|
||||
from customers
|
||||
left join payment_summary using (customer_id)
|
||||
```
|
||||
|
||||
### Pattern 4: Mart Models (Dimensions and Facts)
|
||||
|
||||
```sql
|
||||
-- models/marts/core/dim_customers.sql
|
||||
{{
|
||||
config(
|
||||
materialized='table',
|
||||
unique_key='customer_id'
|
||||
)
|
||||
}}
|
||||
|
||||
with customers as (
|
||||
select * from {{ ref('int_payments_pivoted_to_customer') }}
|
||||
),
|
||||
|
||||
orders as (
|
||||
select * from {{ ref('stg_shopify__orders') }}
|
||||
),
|
||||
|
||||
order_summary as (
|
||||
select
|
||||
customer_id,
|
||||
count(*) as total_orders,
|
||||
sum(total_price) as total_order_value,
|
||||
min(created_at) as first_order_at,
|
||||
max(created_at) as last_order_at
|
||||
from orders
|
||||
group by customer_id
|
||||
),
|
||||
|
||||
final as (
|
||||
select
|
||||
-- surrogate key
|
||||
{{ dbt_utils.generate_surrogate_key(['customers.customer_id']) }} as customer_key,
|
||||
|
||||
-- natural key
|
||||
customers.customer_id,
|
||||
|
||||
-- attributes
|
||||
customers.email,
|
||||
customers.customer_created_at,
|
||||
|
||||
-- payment metrics
|
||||
customers.total_payments,
|
||||
customers.successful_payments,
|
||||
customers.lifetime_value,
|
||||
customers.first_payment_at,
|
||||
customers.last_payment_at,
|
||||
|
||||
-- order metrics
|
||||
coalesce(order_summary.total_orders, 0) as total_orders,
|
||||
coalesce(order_summary.total_order_value, 0) as total_order_value,
|
||||
order_summary.first_order_at,
|
||||
order_summary.last_order_at,
|
||||
|
||||
-- calculated fields
|
||||
case
|
||||
when customers.lifetime_value >= 1000 then 'high'
|
||||
when customers.lifetime_value >= 100 then 'medium'
|
||||
else 'low'
|
||||
end as customer_tier,
|
||||
|
||||
-- timestamps
|
||||
current_timestamp as _loaded_at
|
||||
|
||||
from customers
|
||||
left join order_summary using (customer_id)
|
||||
)
|
||||
|
||||
select * from final
|
||||
```
|
||||
|
||||
```sql
|
||||
-- models/marts/core/fct_orders.sql
|
||||
{{
|
||||
config(
|
||||
materialized='incremental',
|
||||
unique_key='order_id',
|
||||
incremental_strategy='merge'
|
||||
)
|
||||
}}
|
||||
|
||||
with orders as (
|
||||
select * from {{ ref('stg_shopify__orders') }}
|
||||
|
||||
{% if is_incremental() %}
|
||||
where updated_at > (select max(updated_at) from {{ this }})
|
||||
{% endif %}
|
||||
),
|
||||
|
||||
customers as (
|
||||
select * from {{ ref('dim_customers') }}
|
||||
),
|
||||
|
||||
final as (
|
||||
select
|
||||
-- keys
|
||||
orders.order_id,
|
||||
customers.customer_key,
|
||||
orders.customer_id,
|
||||
|
||||
-- dimensions
|
||||
orders.order_status,
|
||||
orders.fulfillment_status,
|
||||
orders.payment_status,
|
||||
|
||||
-- measures
|
||||
orders.subtotal,
|
||||
orders.tax,
|
||||
orders.shipping,
|
||||
orders.total_price,
|
||||
orders.total_discount,
|
||||
orders.item_count,
|
||||
|
||||
-- timestamps
|
||||
orders.created_at,
|
||||
orders.updated_at,
|
||||
orders.fulfilled_at,
|
||||
|
||||
-- metadata
|
||||
current_timestamp as _loaded_at
|
||||
|
||||
from orders
|
||||
left join customers on orders.customer_id = customers.customer_id
|
||||
)
|
||||
|
||||
select * from final
|
||||
```
|
||||
|
||||
### Pattern 5: Testing and Documentation
|
||||
|
||||
```yaml
|
||||
# models/marts/core/_core__models.yml
|
||||
version: 2
|
||||
|
||||
models:
|
||||
- name: dim_customers
|
||||
description: Customer dimension with payment and order metrics
|
||||
columns:
|
||||
- name: customer_key
|
||||
description: Surrogate key for the customer dimension
|
||||
tests:
|
||||
- unique
|
||||
- not_null
|
||||
|
||||
- name: customer_id
|
||||
description: Natural key from source system
|
||||
tests:
|
||||
- unique
|
||||
- not_null
|
||||
|
||||
- name: email
|
||||
description: Customer email address
|
||||
tests:
|
||||
- not_null
|
||||
|
||||
- name: customer_tier
|
||||
description: Customer value tier based on lifetime value
|
||||
tests:
|
||||
- accepted_values:
|
||||
values: ['high', 'medium', 'low']
|
||||
|
||||
- name: lifetime_value
|
||||
description: Total amount paid by customer
|
||||
tests:
|
||||
- dbt_utils.expression_is_true:
|
||||
expression: ">= 0"
|
||||
|
||||
- name: fct_orders
|
||||
description: Order fact table with all order transactions
|
||||
tests:
|
||||
- dbt_utils.recency:
|
||||
datepart: day
|
||||
field: created_at
|
||||
interval: 1
|
||||
columns:
|
||||
- name: order_id
|
||||
tests:
|
||||
- unique
|
||||
- not_null
|
||||
- name: customer_key
|
||||
tests:
|
||||
- not_null
|
||||
- relationships:
|
||||
to: ref('dim_customers')
|
||||
field: customer_key
|
||||
```
|
||||
|
||||
### Pattern 6: Macros and DRY Code
|
||||
|
||||
```sql
|
||||
-- macros/cents_to_dollars.sql
|
||||
{% macro cents_to_dollars(column_name, precision=2) %}
|
||||
round({{ column_name }} / 100.0, {{ precision }})
|
||||
{% endmacro %}
|
||||
|
||||
-- macros/generate_schema_name.sql
|
||||
{% macro generate_schema_name(custom_schema_name, node) %}
|
||||
{%- set default_schema = target.schema -%}
|
||||
{%- if custom_schema_name is none -%}
|
||||
{{ default_schema }}
|
||||
{%- else -%}
|
||||
{{ default_schema }}_{{ custom_schema_name }}
|
||||
{%- endif -%}
|
||||
{% endmacro %}
|
||||
|
||||
-- macros/limit_data_in_dev.sql
|
||||
{% macro limit_data_in_dev(column_name, days=3) %}
|
||||
{% if target.name == 'dev' %}
|
||||
where {{ column_name }} >= dateadd(day, -{{ days }}, current_date)
|
||||
{% endif %}
|
||||
{% endmacro %}
|
||||
|
||||
-- Usage in model
|
||||
select * from {{ ref('stg_orders') }}
|
||||
{{ limit_data_in_dev('created_at') }}
|
||||
```
|
||||
|
||||
### Pattern 7: Incremental Strategies
|
||||
|
||||
```sql
|
||||
-- Delete+Insert (default for most warehouses)
|
||||
{{
|
||||
config(
|
||||
materialized='incremental',
|
||||
unique_key='id',
|
||||
incremental_strategy='delete+insert'
|
||||
)
|
||||
}}
|
||||
|
||||
-- Merge (best for late-arriving data)
|
||||
{{
|
||||
config(
|
||||
materialized='incremental',
|
||||
unique_key='id',
|
||||
incremental_strategy='merge',
|
||||
merge_update_columns=['status', 'amount', 'updated_at']
|
||||
)
|
||||
}}
|
||||
|
||||
-- Insert Overwrite (partition-based)
|
||||
{{
|
||||
config(
|
||||
materialized='incremental',
|
||||
incremental_strategy='insert_overwrite',
|
||||
partition_by={
|
||||
"field": "created_date",
|
||||
"data_type": "date",
|
||||
"granularity": "day"
|
||||
}
|
||||
)
|
||||
}}
|
||||
|
||||
select
|
||||
*,
|
||||
date(created_at) as created_date
|
||||
from {{ ref('stg_events') }}
|
||||
|
||||
{% if is_incremental() %}
|
||||
where created_date >= dateadd(day, -3, current_date)
|
||||
{% endif %}
|
||||
```
|
||||
|
||||
## dbt Commands
|
||||
|
||||
```bash
|
||||
# Development
|
||||
dbt run # Run all models
|
||||
dbt run --select staging # Run staging models only
|
||||
dbt run --select +fct_orders # Run fct_orders and its upstream
|
||||
dbt run --select fct_orders+ # Run fct_orders and its downstream
|
||||
dbt run --full-refresh # Rebuild incremental models
|
||||
|
||||
# Testing
|
||||
dbt test # Run all tests
|
||||
dbt test --select stg_stripe # Test specific models
|
||||
dbt build # Run + test in DAG order
|
||||
|
||||
# Documentation
|
||||
dbt docs generate # Generate docs
|
||||
dbt docs serve # Serve docs locally
|
||||
|
||||
# Debugging
|
||||
dbt compile # Compile SQL without running
|
||||
dbt debug # Test connection
|
||||
dbt ls --select tag:critical # List models by tag
|
||||
```
|
||||
|
||||
## Best Practices
|
||||
|
||||
### Do's
|
||||
- **Use staging layer** - Clean data once, use everywhere
|
||||
- **Test aggressively** - Not null, unique, relationships
|
||||
- **Document everything** - Column descriptions, model descriptions
|
||||
- **Use incremental** - For tables > 1M rows
|
||||
- **Version control** - dbt project in Git
|
||||
|
||||
### Don'ts
|
||||
- **Don't skip staging** - Raw → mart is tech debt
|
||||
- **Don't hardcode dates** - Use `{{ var('start_date') }}`
|
||||
- **Don't repeat logic** - Extract to macros
|
||||
- **Don't test in prod** - Use dev target
|
||||
- **Don't ignore freshness** - Monitor source data
|
||||
|
||||
## Resources
|
||||
|
||||
- [dbt Documentation](https://docs.getdbt.com/)
|
||||
- [dbt Best Practices](https://docs.getdbt.com/guides/best-practices)
|
||||
- [dbt-utils Package](https://hub.getdbt.com/dbt-labs/dbt_utils/latest/)
|
||||
- [dbt Discourse](https://discourse.getdbt.com/)
|
||||
+499
@@ -0,0 +1,499 @@
|
||||
---
|
||||
name: embedding-strategies
|
||||
description: "Guide to selecting and optimizing embedding models for vector search applications."
|
||||
risk: unknown
|
||||
source: community
|
||||
date_added: "2026-02-27"
|
||||
---
|
||||
|
||||
# Embedding Strategies
|
||||
|
||||
Guide to selecting and optimizing embedding models for vector search applications.
|
||||
|
||||
## Do not use this skill when
|
||||
|
||||
- The task is unrelated to embedding strategies
|
||||
- You need a different domain or tool outside this scope
|
||||
|
||||
## Instructions
|
||||
|
||||
- Clarify goals, constraints, and required inputs.
|
||||
- Apply relevant best practices and validate outcomes.
|
||||
- Provide actionable steps and verification.
|
||||
- If detailed examples are required, open `resources/implementation-playbook.md`.
|
||||
|
||||
## Use this skill when
|
||||
|
||||
- Choosing embedding models for RAG
|
||||
- Optimizing chunking strategies
|
||||
- Fine-tuning embeddings for domains
|
||||
- Comparing embedding model performance
|
||||
- Reducing embedding dimensions
|
||||
- Handling multilingual content
|
||||
|
||||
## Core Concepts
|
||||
|
||||
### 1. Embedding Model Comparison
|
||||
|
||||
| Model | Dimensions | Max Tokens | Best For |
|
||||
|-------|------------|------------|----------|
|
||||
| **text-embedding-3-large** | 3072 | 8191 | High accuracy |
|
||||
| **text-embedding-3-small** | 1536 | 8191 | Cost-effective |
|
||||
| **voyage-2** | 1024 | 4000 | Code, legal |
|
||||
| **bge-large-en-v1.5** | 1024 | 512 | Open source |
|
||||
| **all-MiniLM-L6-v2** | 384 | 256 | Fast, lightweight |
|
||||
| **multilingual-e5-large** | 1024 | 512 | Multi-language |
|
||||
|
||||
### 2. Embedding Pipeline
|
||||
|
||||
```
|
||||
Document → Chunking → Preprocessing → Embedding Model → Vector
|
||||
↓
|
||||
[Overlap, Size] [Clean, Normalize] [API/Local]
|
||||
```
|
||||
|
||||
## Templates
|
||||
|
||||
### Template 1: OpenAI Embeddings
|
||||
|
||||
```python
|
||||
from openai import OpenAI
|
||||
from typing import List
|
||||
import numpy as np
|
||||
|
||||
client = OpenAI()
|
||||
|
||||
def get_embeddings(
|
||||
texts: List[str],
|
||||
model: str = "text-embedding-3-small",
|
||||
dimensions: int = None
|
||||
) -> List[List[float]]:
|
||||
"""Get embeddings from OpenAI."""
|
||||
# Handle batching for large lists
|
||||
batch_size = 100
|
||||
all_embeddings = []
|
||||
|
||||
for i in range(0, len(texts), batch_size):
|
||||
batch = texts[i:i + batch_size]
|
||||
|
||||
kwargs = {"input": batch, "model": model}
|
||||
if dimensions:
|
||||
kwargs["dimensions"] = dimensions
|
||||
|
||||
response = client.embeddings.create(**kwargs)
|
||||
embeddings = [item.embedding for item in response.data]
|
||||
all_embeddings.extend(embeddings)
|
||||
|
||||
return all_embeddings
|
||||
|
||||
|
||||
def get_embedding(text: str, **kwargs) -> List[float]:
|
||||
"""Get single embedding."""
|
||||
return get_embeddings([text], **kwargs)[0]
|
||||
|
||||
|
||||
# Dimension reduction with OpenAI
|
||||
def get_reduced_embedding(text: str, dimensions: int = 512) -> List[float]:
|
||||
"""Get embedding with reduced dimensions (Matryoshka)."""
|
||||
return get_embedding(
|
||||
text,
|
||||
model="text-embedding-3-small",
|
||||
dimensions=dimensions
|
||||
)
|
||||
```
|
||||
|
||||
### Template 2: Local Embeddings with Sentence Transformers
|
||||
|
||||
```python
|
||||
from sentence_transformers import SentenceTransformer
|
||||
from typing import List, Optional
|
||||
import numpy as np
|
||||
|
||||
class LocalEmbedder:
|
||||
"""Local embedding with sentence-transformers."""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
model_name: str = "BAAI/bge-large-en-v1.5",
|
||||
device: str = "cuda"
|
||||
):
|
||||
self.model = SentenceTransformer(model_name, device=device)
|
||||
|
||||
def embed(
|
||||
self,
|
||||
texts: List[str],
|
||||
normalize: bool = True,
|
||||
show_progress: bool = False
|
||||
) -> np.ndarray:
|
||||
"""Embed texts with optional normalization."""
|
||||
embeddings = self.model.encode(
|
||||
texts,
|
||||
normalize_embeddings=normalize,
|
||||
show_progress_bar=show_progress,
|
||||
convert_to_numpy=True
|
||||
)
|
||||
return embeddings
|
||||
|
||||
def embed_query(self, query: str) -> np.ndarray:
|
||||
"""Embed a query with BGE-style prefix."""
|
||||
# BGE models benefit from query prefix
|
||||
if "bge" in self.model.get_sentence_embedding_dimension():
|
||||
query = f"Represent this sentence for searching relevant passages: {query}"
|
||||
return self.embed([query])[0]
|
||||
|
||||
def embed_documents(self, documents: List[str]) -> np.ndarray:
|
||||
"""Embed documents for indexing."""
|
||||
return self.embed(documents)
|
||||
|
||||
|
||||
# E5 model with instructions
|
||||
class E5Embedder:
|
||||
def __init__(self, model_name: str = "intfloat/multilingual-e5-large"):
|
||||
self.model = SentenceTransformer(model_name)
|
||||
|
||||
def embed_query(self, query: str) -> np.ndarray:
|
||||
return self.model.encode(f"query: {query}")
|
||||
|
||||
def embed_document(self, document: str) -> np.ndarray:
|
||||
return self.model.encode(f"passage: {document}")
|
||||
```
|
||||
|
||||
### Template 3: Chunking Strategies
|
||||
|
||||
```python
|
||||
from typing import List, Tuple
|
||||
import re
|
||||
|
||||
def chunk_by_tokens(
|
||||
text: str,
|
||||
chunk_size: int = 512,
|
||||
chunk_overlap: int = 50,
|
||||
tokenizer=None
|
||||
) -> List[str]:
|
||||
"""Chunk text by token count."""
|
||||
import tiktoken
|
||||
tokenizer = tokenizer or tiktoken.get_encoding("cl100k_base")
|
||||
|
||||
tokens = tokenizer.encode(text)
|
||||
chunks = []
|
||||
|
||||
start = 0
|
||||
while start < len(tokens):
|
||||
end = start + chunk_size
|
||||
chunk_tokens = tokens[start:end]
|
||||
chunk_text = tokenizer.decode(chunk_tokens)
|
||||
chunks.append(chunk_text)
|
||||
start = end - chunk_overlap
|
||||
|
||||
return chunks
|
||||
|
||||
|
||||
def chunk_by_sentences(
|
||||
text: str,
|
||||
max_chunk_size: int = 1000,
|
||||
min_chunk_size: int = 100
|
||||
) -> List[str]:
|
||||
"""Chunk text by sentences, respecting size limits."""
|
||||
import nltk
|
||||
sentences = nltk.sent_tokenize(text)
|
||||
|
||||
chunks = []
|
||||
current_chunk = []
|
||||
current_size = 0
|
||||
|
||||
for sentence in sentences:
|
||||
sentence_size = len(sentence)
|
||||
|
||||
if current_size + sentence_size > max_chunk_size and current_chunk:
|
||||
chunks.append(" ".join(current_chunk))
|
||||
current_chunk = []
|
||||
current_size = 0
|
||||
|
||||
current_chunk.append(sentence)
|
||||
current_size += sentence_size
|
||||
|
||||
if current_chunk:
|
||||
chunks.append(" ".join(current_chunk))
|
||||
|
||||
return chunks
|
||||
|
||||
|
||||
def chunk_by_semantic_sections(
|
||||
text: str,
|
||||
headers_pattern: str = r'^#{1,3}\s+.+$'
|
||||
) -> List[Tuple[str, str]]:
|
||||
"""Chunk markdown by headers, preserving hierarchy."""
|
||||
lines = text.split('\n')
|
||||
chunks = []
|
||||
current_header = ""
|
||||
current_content = []
|
||||
|
||||
for line in lines:
|
||||
if re.match(headers_pattern, line, re.MULTILINE):
|
||||
if current_content:
|
||||
chunks.append((current_header, '\n'.join(current_content)))
|
||||
current_header = line
|
||||
current_content = []
|
||||
else:
|
||||
current_content.append(line)
|
||||
|
||||
if current_content:
|
||||
chunks.append((current_header, '\n'.join(current_content)))
|
||||
|
||||
return chunks
|
||||
|
||||
|
||||
def recursive_character_splitter(
|
||||
text: str,
|
||||
chunk_size: int = 1000,
|
||||
chunk_overlap: int = 200,
|
||||
separators: List[str] = None
|
||||
) -> List[str]:
|
||||
"""LangChain-style recursive splitter."""
|
||||
separators = separators or ["\n\n", "\n", ". ", " ", ""]
|
||||
|
||||
def split_text(text: str, separators: List[str]) -> List[str]:
|
||||
if not text:
|
||||
return []
|
||||
|
||||
separator = separators[0]
|
||||
remaining_separators = separators[1:]
|
||||
|
||||
if separator == "":
|
||||
# Character-level split
|
||||
return [text[i:i+chunk_size] for i in range(0, len(text), chunk_size - chunk_overlap)]
|
||||
|
||||
splits = text.split(separator)
|
||||
chunks = []
|
||||
current_chunk = []
|
||||
current_length = 0
|
||||
|
||||
for split in splits:
|
||||
split_length = len(split) + len(separator)
|
||||
|
||||
if current_length + split_length > chunk_size and current_chunk:
|
||||
chunk_text = separator.join(current_chunk)
|
||||
|
||||
# Recursively split if still too large
|
||||
if len(chunk_text) > chunk_size and remaining_separators:
|
||||
chunks.extend(split_text(chunk_text, remaining_separators))
|
||||
else:
|
||||
chunks.append(chunk_text)
|
||||
|
||||
# Start new chunk with overlap
|
||||
overlap_splits = []
|
||||
overlap_length = 0
|
||||
for s in reversed(current_chunk):
|
||||
if overlap_length + len(s) <= chunk_overlap:
|
||||
overlap_splits.insert(0, s)
|
||||
overlap_length += len(s)
|
||||
else:
|
||||
break
|
||||
current_chunk = overlap_splits
|
||||
current_length = overlap_length
|
||||
|
||||
current_chunk.append(split)
|
||||
current_length += split_length
|
||||
|
||||
if current_chunk:
|
||||
chunks.append(separator.join(current_chunk))
|
||||
|
||||
return chunks
|
||||
|
||||
return split_text(text, separators)
|
||||
```
|
||||
|
||||
### Template 4: Domain-Specific Embedding Pipeline
|
||||
|
||||
```python
|
||||
class DomainEmbeddingPipeline:
|
||||
"""Pipeline for domain-specific embeddings."""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
embedding_model: str = "text-embedding-3-small",
|
||||
chunk_size: int = 512,
|
||||
chunk_overlap: int = 50,
|
||||
preprocessing_fn=None
|
||||
):
|
||||
self.embedding_model = embedding_model
|
||||
self.chunk_size = chunk_size
|
||||
self.chunk_overlap = chunk_overlap
|
||||
self.preprocess = preprocessing_fn or self._default_preprocess
|
||||
|
||||
def _default_preprocess(self, text: str) -> str:
|
||||
"""Default preprocessing."""
|
||||
# Remove excessive whitespace
|
||||
text = re.sub(r'\s+', ' ', text)
|
||||
# Remove special characters
|
||||
text = re.sub(r'[^\w\s.,!?-]', '', text)
|
||||
return text.strip()
|
||||
|
||||
async def process_documents(
|
||||
self,
|
||||
documents: List[dict],
|
||||
id_field: str = "id",
|
||||
content_field: str = "content",
|
||||
metadata_fields: List[str] = None
|
||||
) -> List[dict]:
|
||||
"""Process documents for vector storage."""
|
||||
processed = []
|
||||
|
||||
for doc in documents:
|
||||
content = doc[content_field]
|
||||
doc_id = doc[id_field]
|
||||
|
||||
# Preprocess
|
||||
cleaned = self.preprocess(content)
|
||||
|
||||
# Chunk
|
||||
chunks = chunk_by_tokens(
|
||||
cleaned,
|
||||
self.chunk_size,
|
||||
self.chunk_overlap
|
||||
)
|
||||
|
||||
# Create embeddings
|
||||
embeddings = get_embeddings(chunks, self.embedding_model)
|
||||
|
||||
# Create records
|
||||
for i, (chunk, embedding) in enumerate(zip(chunks, embeddings)):
|
||||
record = {
|
||||
"id": f"{doc_id}_chunk_{i}",
|
||||
"document_id": doc_id,
|
||||
"chunk_index": i,
|
||||
"text": chunk,
|
||||
"embedding": embedding
|
||||
}
|
||||
|
||||
# Add metadata
|
||||
if metadata_fields:
|
||||
for field in metadata_fields:
|
||||
if field in doc:
|
||||
record[field] = doc[field]
|
||||
|
||||
processed.append(record)
|
||||
|
||||
return processed
|
||||
|
||||
|
||||
# Code-specific pipeline
|
||||
class CodeEmbeddingPipeline:
|
||||
"""Specialized pipeline for code embeddings."""
|
||||
|
||||
def __init__(self, model: str = "voyage-code-2"):
|
||||
self.model = model
|
||||
|
||||
def chunk_code(self, code: str, language: str) -> List[dict]:
|
||||
"""Chunk code by functions/classes."""
|
||||
import tree_sitter
|
||||
|
||||
# Parse with tree-sitter
|
||||
# Extract functions, classes, methods
|
||||
# Return chunks with context
|
||||
pass
|
||||
|
||||
def embed_with_context(self, chunk: str, context: str) -> List[float]:
|
||||
"""Embed code with surrounding context."""
|
||||
combined = f"Context: {context}\n\nCode:\n{chunk}"
|
||||
return get_embedding(combined, model=self.model)
|
||||
```
|
||||
|
||||
### Template 5: Embedding Quality Evaluation
|
||||
|
||||
```python
|
||||
import numpy as np
|
||||
from typing import List, Tuple
|
||||
|
||||
def evaluate_retrieval_quality(
|
||||
queries: List[str],
|
||||
relevant_docs: List[List[str]], # List of relevant doc IDs per query
|
||||
retrieved_docs: List[List[str]], # List of retrieved doc IDs per query
|
||||
k: int = 10
|
||||
) -> dict:
|
||||
"""Evaluate embedding quality for retrieval."""
|
||||
|
||||
def precision_at_k(relevant: set, retrieved: List[str], k: int) -> float:
|
||||
retrieved_k = retrieved[:k]
|
||||
relevant_retrieved = len(set(retrieved_k) & relevant)
|
||||
return relevant_retrieved / k
|
||||
|
||||
def recall_at_k(relevant: set, retrieved: List[str], k: int) -> float:
|
||||
retrieved_k = retrieved[:k]
|
||||
relevant_retrieved = len(set(retrieved_k) & relevant)
|
||||
return relevant_retrieved / len(relevant) if relevant else 0
|
||||
|
||||
def mrr(relevant: set, retrieved: List[str]) -> float:
|
||||
for i, doc in enumerate(retrieved):
|
||||
if doc in relevant:
|
||||
return 1 / (i + 1)
|
||||
return 0
|
||||
|
||||
def ndcg_at_k(relevant: set, retrieved: List[str], k: int) -> float:
|
||||
dcg = sum(
|
||||
1 / np.log2(i + 2) if doc in relevant else 0
|
||||
for i, doc in enumerate(retrieved[:k])
|
||||
)
|
||||
ideal_dcg = sum(1 / np.log2(i + 2) for i in range(min(len(relevant), k)))
|
||||
return dcg / ideal_dcg if ideal_dcg > 0 else 0
|
||||
|
||||
metrics = {
|
||||
f"precision@{k}": [],
|
||||
f"recall@{k}": [],
|
||||
"mrr": [],
|
||||
f"ndcg@{k}": []
|
||||
}
|
||||
|
||||
for relevant, retrieved in zip(relevant_docs, retrieved_docs):
|
||||
relevant_set = set(relevant)
|
||||
metrics[f"precision@{k}"].append(precision_at_k(relevant_set, retrieved, k))
|
||||
metrics[f"recall@{k}"].append(recall_at_k(relevant_set, retrieved, k))
|
||||
metrics["mrr"].append(mrr(relevant_set, retrieved))
|
||||
metrics[f"ndcg@{k}"].append(ndcg_at_k(relevant_set, retrieved, k))
|
||||
|
||||
return {name: np.mean(values) for name, values in metrics.items()}
|
||||
|
||||
|
||||
def compute_embedding_similarity(
|
||||
embeddings1: np.ndarray,
|
||||
embeddings2: np.ndarray,
|
||||
metric: str = "cosine"
|
||||
) -> np.ndarray:
|
||||
"""Compute similarity matrix between embedding sets."""
|
||||
if metric == "cosine":
|
||||
# Normalize
|
||||
norm1 = embeddings1 / np.linalg.norm(embeddings1, axis=1, keepdims=True)
|
||||
norm2 = embeddings2 / np.linalg.norm(embeddings2, axis=1, keepdims=True)
|
||||
return norm1 @ norm2.T
|
||||
elif metric == "euclidean":
|
||||
from scipy.spatial.distance import cdist
|
||||
return -cdist(embeddings1, embeddings2, metric='euclidean')
|
||||
elif metric == "dot":
|
||||
return embeddings1 @ embeddings2.T
|
||||
```
|
||||
|
||||
## Best Practices
|
||||
|
||||
### Do's
|
||||
- **Match model to use case** - Code vs prose vs multilingual
|
||||
- **Chunk thoughtfully** - Preserve semantic boundaries
|
||||
- **Normalize embeddings** - For cosine similarity
|
||||
- **Batch requests** - More efficient than one-by-one
|
||||
- **Cache embeddings** - Avoid recomputing
|
||||
|
||||
### Don'ts
|
||||
- **Don't ignore token limits** - Truncation loses info
|
||||
- **Don't mix embedding models** - Incompatible spaces
|
||||
- **Don't skip preprocessing** - Garbage in, garbage out
|
||||
- **Don't over-chunk** - Lose context
|
||||
|
||||
## Resources
|
||||
|
||||
- [OpenAI Embeddings](https://platform.openai.com/docs/guides/embeddings)
|
||||
- [Sentence Transformers](https://www.sbert.net/)
|
||||
- [MTEB Benchmark](https://huggingface.co/spaces/mteb/leaderboard)
|
||||
|
||||
## Limitations
|
||||
- Use this skill only when the task clearly matches the scope described above.
|
||||
- Do not treat the output as a substitute for environment-specific validation, testing, or expert review.
|
||||
- Stop and ask for clarification if required inputs, permissions, safety boundaries, or success criteria are missing.
|
||||
+1528
File diff suppressed because it is too large
Load Diff
+1490
File diff suppressed because it is too large
Load Diff
+119
@@ -0,0 +1,119 @@
|
||||
# Postgres Best Practices - Contributor Guide
|
||||
|
||||
This repository contains Postgres performance optimization rules optimized for
|
||||
AI agents and LLMs.
|
||||
|
||||
## Quick Start
|
||||
|
||||
```bash
|
||||
# Install dependencies
|
||||
cd packages/postgres-best-practices-build
|
||||
npm install
|
||||
|
||||
# Validate existing rules
|
||||
npm run validate
|
||||
|
||||
# Build AGENTS.md
|
||||
npm run build
|
||||
```
|
||||
|
||||
## Creating a New Rule
|
||||
|
||||
1. **Choose a section prefix** based on the category:
|
||||
- `query-` Query Performance (CRITICAL)
|
||||
- `conn-` Connection Management (CRITICAL)
|
||||
- `security-` Security & RLS (CRITICAL)
|
||||
- `schema-` Schema Design (HIGH)
|
||||
- `lock-` Concurrency & Locking (MEDIUM-HIGH)
|
||||
- `data-` Data Access Patterns (MEDIUM)
|
||||
- `monitor-` Monitoring & Diagnostics (LOW-MEDIUM)
|
||||
- `advanced-` Advanced Features (LOW)
|
||||
|
||||
2. **Copy the template**:
|
||||
```bash
|
||||
cp rules/_template.md rules/query-your-rule-name.md
|
||||
```
|
||||
|
||||
3. **Fill in the content** following the template structure
|
||||
|
||||
4. **Validate and build**:
|
||||
```bash
|
||||
npm run validate
|
||||
npm run build
|
||||
```
|
||||
|
||||
5. **Review** the generated `AGENTS.md`
|
||||
|
||||
## Repository Structure
|
||||
|
||||
```
|
||||
skills/postgres-best-practices/
|
||||
├── SKILL.md # Agent-facing skill manifest
|
||||
├── AGENTS.md # [GENERATED] Compiled rules document
|
||||
├── README.md # This file
|
||||
├── metadata.json # Version and metadata
|
||||
└── rules/
|
||||
├── _template.md # Rule template
|
||||
├── _sections.md # Section definitions
|
||||
├── _contributing.md # Writing guidelines
|
||||
└── *.md # Individual rules
|
||||
|
||||
packages/postgres-best-practices-build/
|
||||
├── src/ # Build system source
|
||||
├── package.json # NPM scripts
|
||||
└── test-cases.json # [GENERATED] Test artifacts
|
||||
```
|
||||
|
||||
## Rule File Structure
|
||||
|
||||
See `rules/_template.md` for the complete template. Key elements:
|
||||
|
||||
````markdown
|
||||
---
|
||||
title: Clear, Action-Oriented Title
|
||||
impact: CRITICAL|HIGH|MEDIUM-HIGH|MEDIUM|LOW-MEDIUM|LOW
|
||||
impactDescription: Quantified benefit (e.g., "10-100x faster")
|
||||
tags: relevant, keywords
|
||||
---
|
||||
|
||||
## [Title]
|
||||
|
||||
[1-2 sentence explanation]
|
||||
|
||||
**Incorrect (description):**
|
||||
|
||||
```sql
|
||||
-- Comment explaining what's wrong
|
||||
[Bad SQL example]
|
||||
```
|
||||
````
|
||||
|
||||
**Correct (description):**
|
||||
|
||||
```sql
|
||||
-- Comment explaining why this is better
|
||||
[Good SQL example]
|
||||
```
|
||||
|
||||
```
|
||||
## Writing Guidelines
|
||||
|
||||
See `rules/_contributing.md` for detailed guidelines. Key principles:
|
||||
|
||||
1. **Show concrete transformations** - "Change X to Y", not abstract advice
|
||||
2. **Error-first structure** - Show the problem before the solution
|
||||
3. **Quantify impact** - Include specific metrics (10x faster, 50% smaller)
|
||||
4. **Self-contained examples** - Complete, runnable SQL
|
||||
5. **Semantic naming** - Use meaningful names (users, email), not (table1, col1)
|
||||
|
||||
## Impact Levels
|
||||
|
||||
| Level | Improvement | Examples |
|
||||
|-------|-------------|----------|
|
||||
| CRITICAL | 10-100x | Missing indexes, connection exhaustion |
|
||||
| HIGH | 5-20x | Wrong index types, poor partitioning |
|
||||
| MEDIUM-HIGH | 2-5x | N+1 queries, RLS optimization |
|
||||
| MEDIUM | 1.5-3x | Redundant indexes, stale statistics |
|
||||
| LOW-MEDIUM | 1.2-2x | VACUUM tuning, config tweaks |
|
||||
| LOW | Incremental | Advanced patterns, edge cases |
|
||||
```
|
||||
+63
@@ -0,0 +1,63 @@
|
||||
---
|
||||
name: postgres-best-practices
|
||||
description: "Postgres performance optimization and best practices from Supabase. Use this skill when writing, reviewing, or optimizing Postgres queries, schema designs, or database configurations."
|
||||
risk: safe
|
||||
source: community
|
||||
date_added: "2026-02-27"
|
||||
---
|
||||
|
||||
# Supabase Postgres Best Practices
|
||||
|
||||
Comprehensive performance optimization guide for Postgres, maintained by Supabase. Contains rules across 8 categories, prioritized by impact to guide automated query optimization and schema design.
|
||||
|
||||
## When to Use
|
||||
Reference these guidelines when:
|
||||
- Writing SQL queries or designing schemas
|
||||
- Implementing indexes or query optimization
|
||||
- Reviewing database performance issues
|
||||
- Configuring connection pooling or scaling
|
||||
- Optimizing for Postgres-specific features
|
||||
- Working with Row-Level Security (RLS)
|
||||
|
||||
## Rule Categories by Priority
|
||||
|
||||
| Priority | Category | Impact | Prefix |
|
||||
|----------|----------|--------|--------|
|
||||
| 1 | Query Performance | CRITICAL | `query-` |
|
||||
| 2 | Connection Management | CRITICAL | `conn-` |
|
||||
| 3 | Security & RLS | CRITICAL | `security-` |
|
||||
| 4 | Schema Design | HIGH | `schema-` |
|
||||
| 5 | Concurrency & Locking | MEDIUM-HIGH | `lock-` |
|
||||
| 6 | Data Access Patterns | MEDIUM | `data-` |
|
||||
| 7 | Monitoring & Diagnostics | LOW-MEDIUM | `monitor-` |
|
||||
| 8 | Advanced Features | LOW | `advanced-` |
|
||||
|
||||
## How to Use
|
||||
|
||||
Read individual rule files for detailed explanations and SQL examples:
|
||||
|
||||
```
|
||||
rules/query-missing-indexes.md
|
||||
rules/schema-partial-indexes.md
|
||||
rules/_sections.md
|
||||
```
|
||||
|
||||
Each rule file contains:
|
||||
- Brief explanation of why it matters
|
||||
- Incorrect SQL example with explanation
|
||||
- Correct SQL example with explanation
|
||||
- Optional EXPLAIN output or metrics
|
||||
- Additional context and references
|
||||
- Supabase-specific notes (when applicable)
|
||||
|
||||
## Full Compiled Document
|
||||
|
||||
For the complete guide with all rules expanded: `AGENTS.md`
|
||||
|
||||
### When to Use
|
||||
This skill is applicable to execute the workflow or actions described in the overview.
|
||||
|
||||
## Limitations
|
||||
- Use this skill only when the task clearly matches the scope described above.
|
||||
- Do not treat the output as a substitute for environment-specific validation, testing, or expert review.
|
||||
- Stop and ask for clarification if required inputs, permissions, safety boundaries, or success criteria are missing.
|
||||
+13
@@ -0,0 +1,13 @@
|
||||
{
|
||||
"version": "1.0.0",
|
||||
"organization": "Supabase",
|
||||
"date": "January 2026",
|
||||
"abstract": "Comprehensive Postgres performance optimization guide for developers using Supabase and Postgres. Contains performance rules across 8 categories, prioritized by impact from critical (query performance, connection management) to incremental (advanced features). Each rule includes detailed explanations, incorrect vs. correct SQL examples, query plan analysis, and specific performance metrics to guide automated optimization and code generation.",
|
||||
"references": [
|
||||
"https://www.postgresql.org/docs/current/",
|
||||
"https://supabase.com/docs",
|
||||
"https://wiki.postgresql.org/wiki/Performance_Optimization",
|
||||
"https://supabase.com/docs/guides/database/overview",
|
||||
"https://supabase.com/docs/guides/auth/row-level-security"
|
||||
]
|
||||
}
|
||||
+171
@@ -0,0 +1,171 @@
|
||||
# Writing Guidelines for Postgres Rules
|
||||
|
||||
This document provides guidelines for creating effective Postgres best
|
||||
practice rules that work well with AI agents and LLMs.
|
||||
|
||||
## Key Principles
|
||||
|
||||
### 1. Concrete Transformation Patterns
|
||||
|
||||
Show exact SQL rewrites. Avoid philosophical advice.
|
||||
|
||||
**Good:** "Use `WHERE id = ANY(ARRAY[...])` instead of
|
||||
`WHERE id IN (SELECT ...)`" **Bad:** "Design good schemas"
|
||||
|
||||
### 2. Error-First Structure
|
||||
|
||||
Always show the problematic pattern first, then the solution. This trains agents
|
||||
to recognize anti-patterns.
|
||||
|
||||
```markdown
|
||||
**Incorrect (sequential queries):** [bad example]
|
||||
|
||||
**Correct (batched query):** [good example]
|
||||
```
|
||||
|
||||
### 3. Quantified Impact
|
||||
|
||||
Include specific metrics. Helps agents prioritize fixes.
|
||||
|
||||
**Good:** "10x faster queries", "50% smaller index", "Eliminates N+1"
|
||||
**Bad:** "Faster", "Better", "More efficient"
|
||||
|
||||
### 4. Self-Contained Examples
|
||||
|
||||
Examples should be complete and runnable (or close to it). Include `CREATE TABLE`
|
||||
if context is needed.
|
||||
|
||||
```sql
|
||||
-- Include table definition when needed for clarity
|
||||
CREATE TABLE users (
|
||||
id bigint PRIMARY KEY,
|
||||
email text NOT NULL,
|
||||
deleted_at timestamptz
|
||||
);
|
||||
|
||||
-- Now show the index
|
||||
CREATE INDEX users_active_email_idx ON users(email) WHERE deleted_at IS NULL;
|
||||
```
|
||||
|
||||
### 5. Semantic Naming
|
||||
|
||||
Use meaningful table/column names. Names carry intent for LLMs.
|
||||
|
||||
**Good:** `users`, `email`, `created_at`, `is_active`
|
||||
**Bad:** `table1`, `col1`, `field`, `flag`
|
||||
|
||||
---
|
||||
|
||||
## Code Example Standards
|
||||
|
||||
### SQL Formatting
|
||||
|
||||
```sql
|
||||
-- Use lowercase keywords, clear formatting
|
||||
CREATE INDEX CONCURRENTLY users_email_idx
|
||||
ON users(email)
|
||||
WHERE deleted_at IS NULL;
|
||||
|
||||
-- Not cramped or ALL CAPS
|
||||
CREATE INDEX CONCURRENTLY USERS_EMAIL_IDX ON USERS(EMAIL) WHERE DELETED_AT IS NULL;
|
||||
```
|
||||
|
||||
### Comments
|
||||
|
||||
- Explain _why_, not _what_
|
||||
- Highlight performance implications
|
||||
- Point out common pitfalls
|
||||
|
||||
### Language Tags
|
||||
|
||||
- `sql` - Standard SQL queries
|
||||
- `plpgsql` - Stored procedures/functions
|
||||
- `typescript` - Application code (when needed)
|
||||
- `python` - Application code (when needed)
|
||||
|
||||
---
|
||||
|
||||
## When to Include Application Code
|
||||
|
||||
**Default: SQL Only**
|
||||
|
||||
Most rules should focus on pure SQL patterns. This keeps examples portable.
|
||||
|
||||
**Include Application Code When:**
|
||||
|
||||
- Connection pooling configuration
|
||||
- Transaction management in application context
|
||||
- ORM anti-patterns (N+1 in Prisma/TypeORM)
|
||||
- Prepared statement usage
|
||||
|
||||
**Format for Mixed Examples:**
|
||||
|
||||
````markdown
|
||||
**Incorrect (N+1 in application):**
|
||||
|
||||
```typescript
|
||||
for (const user of users) {
|
||||
const posts = await db.query("SELECT * FROM posts WHERE user_id = $1", [
|
||||
user.id,
|
||||
]);
|
||||
}
|
||||
```
|
||||
````
|
||||
|
||||
**Correct (batch query):**
|
||||
|
||||
```typescript
|
||||
const posts = await db.query("SELECT * FROM posts WHERE user_id = ANY($1)", [
|
||||
userIds,
|
||||
]);
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Impact Level Guidelines
|
||||
|
||||
| Level | Improvement | Use When |
|
||||
|-------|-------------|----------|
|
||||
| **CRITICAL** | 10-100x | Missing indexes, connection exhaustion, sequential scans on large tables |
|
||||
| **HIGH** | 5-20x | Wrong index types, poor partitioning, missing covering indexes |
|
||||
| **MEDIUM-HIGH** | 2-5x | N+1 queries, inefficient pagination, RLS optimization |
|
||||
| **MEDIUM** | 1.5-3x | Redundant indexes, query plan instability |
|
||||
| **LOW-MEDIUM** | 1.2-2x | VACUUM tuning, configuration tweaks |
|
||||
| **LOW** | Incremental | Advanced patterns, edge cases |
|
||||
|
||||
---
|
||||
|
||||
## Reference Standards
|
||||
|
||||
**Primary Sources:**
|
||||
|
||||
- Official Postgres documentation
|
||||
- Supabase documentation
|
||||
- Postgres wiki
|
||||
- Established blogs (2ndQuadrant, Crunchy Data)
|
||||
|
||||
**Format:**
|
||||
|
||||
```markdown
|
||||
Reference:
|
||||
[Postgres Indexes](https://www.postgresql.org/docs/current/indexes.html)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Review Checklist
|
||||
|
||||
Before submitting a rule:
|
||||
|
||||
- [ ] Title is clear and action-oriented
|
||||
- [ ] Impact level matches the performance gain
|
||||
- [ ] impactDescription includes quantification
|
||||
- [ ] Explanation is concise (1-2 sentences)
|
||||
- [ ] Has at least 1 **Incorrect** SQL example
|
||||
- [ ] Has at least 1 **Correct** SQL example
|
||||
- [ ] SQL uses semantic naming
|
||||
- [ ] Comments explain _why_, not _what_
|
||||
- [ ] Trade-offs mentioned if applicable
|
||||
- [ ] Reference links included
|
||||
- [ ] `npm run validate` passes
|
||||
- [ ] `npm run build` generates correct output
|
||||
+39
@@ -0,0 +1,39 @@
|
||||
# Section Definitions
|
||||
|
||||
This file defines the rule categories for Postgres best practices. Rules are automatically assigned to sections based on their filename prefix.
|
||||
|
||||
Take the examples below as pure demonstrative. Replace each section with the actual rule categories for Postgres best practices.
|
||||
|
||||
---
|
||||
|
||||
## 1. Query Performance (query)
|
||||
**Impact:** CRITICAL
|
||||
**Description:** Slow queries, missing indexes, inefficient query plans. The most common source of Postgres performance issues.
|
||||
|
||||
## 2. Connection Management (conn)
|
||||
**Impact:** CRITICAL
|
||||
**Description:** Connection pooling, limits, and serverless strategies. Critical for applications with high concurrency or serverless deployments.
|
||||
|
||||
## 3. Security & RLS (security)
|
||||
**Impact:** CRITICAL
|
||||
**Description:** Row-Level Security policies, privilege management, and authentication patterns.
|
||||
|
||||
## 4. Schema Design (schema)
|
||||
**Impact:** HIGH
|
||||
**Description:** Table design, index strategies, partitioning, and data type selection. Foundation for long-term performance.
|
||||
|
||||
## 5. Concurrency & Locking (lock)
|
||||
**Impact:** MEDIUM-HIGH
|
||||
**Description:** Transaction management, isolation levels, deadlock prevention, and lock contention patterns.
|
||||
|
||||
## 6. Data Access Patterns (data)
|
||||
**Impact:** MEDIUM
|
||||
**Description:** N+1 query elimination, batch operations, cursor-based pagination, and efficient data fetching.
|
||||
|
||||
## 7. Monitoring & Diagnostics (monitor)
|
||||
**Impact:** LOW-MEDIUM
|
||||
**Description:** Using pg_stat_statements, EXPLAIN ANALYZE, metrics collection, and performance diagnostics.
|
||||
|
||||
## 8. Advanced Features (advanced)
|
||||
**Impact:** LOW
|
||||
**Description:** Full-text search, JSONB optimization, PostGIS, extensions, and advanced Postgres features.
|
||||
+34
@@ -0,0 +1,34 @@
|
||||
---
|
||||
title: Clear, Action-Oriented Title (e.g., "Use Partial Indexes for Filtered Queries")
|
||||
impact: MEDIUM
|
||||
impactDescription: 5-20x query speedup for filtered queries
|
||||
tags: indexes, query-optimization, performance
|
||||
---
|
||||
|
||||
## [Rule Title]
|
||||
|
||||
[1-2 sentence explanation of the problem and why it matters. Focus on performance impact.]
|
||||
|
||||
**Incorrect (describe the problem):**
|
||||
|
||||
```sql
|
||||
-- Comment explaining what makes this slow/problematic
|
||||
CREATE INDEX users_email_idx ON users(email);
|
||||
|
||||
SELECT * FROM users WHERE email = 'user@example.com' AND deleted_at IS NULL;
|
||||
-- This scans deleted records unnecessarily
|
||||
```
|
||||
|
||||
**Correct (describe the solution):**
|
||||
|
||||
```sql
|
||||
-- Comment explaining why this is better
|
||||
CREATE INDEX users_active_email_idx ON users(email) WHERE deleted_at IS NULL;
|
||||
|
||||
SELECT * FROM users WHERE email = 'user@example.com' AND deleted_at IS NULL;
|
||||
-- Only indexes active users, 10x smaller index, faster queries
|
||||
```
|
||||
|
||||
[Optional: Additional context, edge cases, or trade-offs]
|
||||
|
||||
Reference: [Postgres Docs](https://www.postgresql.org/docs/current/)
|
||||
+55
@@ -0,0 +1,55 @@
|
||||
---
|
||||
title: Use tsvector for Full-Text Search
|
||||
impact: MEDIUM
|
||||
impactDescription: 100x faster than LIKE, with ranking support
|
||||
tags: full-text-search, tsvector, gin, search
|
||||
---
|
||||
|
||||
## Use tsvector for Full-Text Search
|
||||
|
||||
LIKE with wildcards can't use indexes. Full-text search with tsvector is orders of magnitude faster.
|
||||
|
||||
**Incorrect (LIKE pattern matching):**
|
||||
|
||||
```sql
|
||||
-- Cannot use index, scans all rows
|
||||
select * from articles where content like '%postgresql%';
|
||||
|
||||
-- Case-insensitive makes it worse
|
||||
select * from articles where lower(content) like '%postgresql%';
|
||||
```
|
||||
|
||||
**Correct (full-text search with tsvector):**
|
||||
|
||||
```sql
|
||||
-- Add tsvector column and index
|
||||
alter table articles add column search_vector tsvector
|
||||
generated always as (to_tsvector('english', coalesce(title,'') || ' ' || coalesce(content,''))) stored;
|
||||
|
||||
create index articles_search_idx on articles using gin (search_vector);
|
||||
|
||||
-- Fast full-text search
|
||||
select * from articles
|
||||
where search_vector @@ to_tsquery('english', 'postgresql & performance');
|
||||
|
||||
-- With ranking
|
||||
select *, ts_rank(search_vector, query) as rank
|
||||
from articles, to_tsquery('english', 'postgresql') query
|
||||
where search_vector @@ query
|
||||
order by rank desc;
|
||||
```
|
||||
|
||||
Search multiple terms:
|
||||
|
||||
```sql
|
||||
-- AND: both terms required
|
||||
to_tsquery('postgresql & performance')
|
||||
|
||||
-- OR: either term
|
||||
to_tsquery('postgresql | mysql')
|
||||
|
||||
-- Prefix matching
|
||||
to_tsquery('post:*')
|
||||
```
|
||||
|
||||
Reference: [Full Text Search](https://supabase.com/docs/guides/database/full-text-search)
|
||||
+49
@@ -0,0 +1,49 @@
|
||||
---
|
||||
title: Index JSONB Columns for Efficient Querying
|
||||
impact: MEDIUM
|
||||
impactDescription: 10-100x faster JSONB queries with proper indexing
|
||||
tags: jsonb, gin, indexes, json
|
||||
---
|
||||
|
||||
## Index JSONB Columns for Efficient Querying
|
||||
|
||||
JSONB queries without indexes scan the entire table. Use GIN indexes for containment queries.
|
||||
|
||||
**Incorrect (no index on JSONB):**
|
||||
|
||||
```sql
|
||||
create table products (
|
||||
id bigint primary key,
|
||||
attributes jsonb
|
||||
);
|
||||
|
||||
-- Full table scan for every query
|
||||
select * from products where attributes @> '{"color": "red"}';
|
||||
select * from products where attributes->>'brand' = 'Nike';
|
||||
```
|
||||
|
||||
**Correct (GIN index for JSONB):**
|
||||
|
||||
```sql
|
||||
-- GIN index for containment operators (@>, ?, ?&, ?|)
|
||||
create index products_attrs_gin on products using gin (attributes);
|
||||
|
||||
-- Now containment queries use the index
|
||||
select * from products where attributes @> '{"color": "red"}';
|
||||
|
||||
-- For specific key lookups, use expression index
|
||||
create index products_brand_idx on products ((attributes->>'brand'));
|
||||
select * from products where attributes->>'brand' = 'Nike';
|
||||
```
|
||||
|
||||
Choose the right operator class:
|
||||
|
||||
```sql
|
||||
-- jsonb_ops (default): supports all operators, larger index
|
||||
create index idx1 on products using gin (attributes);
|
||||
|
||||
-- jsonb_path_ops: only @> operator, but 2-3x smaller index
|
||||
create index idx2 on products using gin (attributes jsonb_path_ops);
|
||||
```
|
||||
|
||||
Reference: [JSONB Indexes](https://www.postgresql.org/docs/current/datatype-json.html#JSON-INDEXING)
|
||||
+46
@@ -0,0 +1,46 @@
|
||||
---
|
||||
title: Configure Idle Connection Timeouts
|
||||
impact: HIGH
|
||||
impactDescription: Reclaim 30-50% of connection slots from idle clients
|
||||
tags: connections, timeout, idle, resource-management
|
||||
---
|
||||
|
||||
## Configure Idle Connection Timeouts
|
||||
|
||||
Idle connections waste resources. Configure timeouts to automatically reclaim them.
|
||||
|
||||
**Incorrect (connections held indefinitely):**
|
||||
|
||||
```sql
|
||||
-- No timeout configured
|
||||
show idle_in_transaction_session_timeout; -- 0 (disabled)
|
||||
|
||||
-- Connections stay open forever, even when idle
|
||||
select pid, state, state_change, query
|
||||
from pg_stat_activity
|
||||
where state = 'idle in transaction';
|
||||
-- Shows transactions idle for hours, holding locks
|
||||
```
|
||||
|
||||
**Correct (automatic cleanup of idle connections):**
|
||||
|
||||
```sql
|
||||
-- Terminate connections idle in transaction after 30 seconds
|
||||
alter system set idle_in_transaction_session_timeout = '30s';
|
||||
|
||||
-- Terminate completely idle connections after 10 minutes
|
||||
alter system set idle_session_timeout = '10min';
|
||||
|
||||
-- Reload configuration
|
||||
select pg_reload_conf();
|
||||
```
|
||||
|
||||
For pooled connections, configure at the pooler level:
|
||||
|
||||
```ini
|
||||
# pgbouncer.ini
|
||||
server_idle_timeout = 60
|
||||
client_idle_timeout = 300
|
||||
```
|
||||
|
||||
Reference: [Connection Timeouts](https://www.postgresql.org/docs/current/runtime-config-client.html#GUC-IDLE-IN-TRANSACTION-SESSION-TIMEOUT)
|
||||
+44
@@ -0,0 +1,44 @@
|
||||
---
|
||||
title: Set Appropriate Connection Limits
|
||||
impact: CRITICAL
|
||||
impactDescription: Prevent database crashes and memory exhaustion
|
||||
tags: connections, max-connections, limits, stability
|
||||
---
|
||||
|
||||
## Set Appropriate Connection Limits
|
||||
|
||||
Too many connections exhaust memory and degrade performance. Set limits based on available resources.
|
||||
|
||||
**Incorrect (unlimited or excessive connections):**
|
||||
|
||||
```sql
|
||||
-- Default max_connections = 100, but often increased blindly
|
||||
show max_connections; -- 500 (way too high for 4GB RAM)
|
||||
|
||||
-- Each connection uses 1-3MB RAM
|
||||
-- 500 connections * 2MB = 1GB just for connections!
|
||||
-- Out of memory errors under load
|
||||
```
|
||||
|
||||
**Correct (calculate based on resources):**
|
||||
|
||||
```sql
|
||||
-- Formula: max_connections = (RAM in MB / 5MB per connection) - reserved
|
||||
-- For 4GB RAM: (4096 / 5) - 10 = ~800 theoretical max
|
||||
-- But practically, 100-200 is better for query performance
|
||||
|
||||
-- Recommended settings for 4GB RAM
|
||||
alter system set max_connections = 100;
|
||||
|
||||
-- Also set work_mem appropriately
|
||||
-- work_mem * max_connections should not exceed 25% of RAM
|
||||
alter system set work_mem = '8MB'; -- 8MB * 100 = 800MB max
|
||||
```
|
||||
|
||||
Monitor connection usage:
|
||||
|
||||
```sql
|
||||
select count(*), state from pg_stat_activity group by state;
|
||||
```
|
||||
|
||||
Reference: [Database Connections](https://supabase.com/docs/guides/platform/performance#connection-management)
|
||||
+41
@@ -0,0 +1,41 @@
|
||||
---
|
||||
title: Use Connection Pooling for All Applications
|
||||
impact: CRITICAL
|
||||
impactDescription: Handle 10-100x more concurrent users
|
||||
tags: connection-pooling, pgbouncer, performance, scalability
|
||||
---
|
||||
|
||||
## Use Connection Pooling for All Applications
|
||||
|
||||
Postgres connections are expensive (1-3MB RAM each). Without pooling, applications exhaust connections under load.
|
||||
|
||||
**Incorrect (new connection per request):**
|
||||
|
||||
```sql
|
||||
-- Each request creates a new connection
|
||||
-- Application code: db.connect() per request
|
||||
-- Result: 500 concurrent users = 500 connections = crashed database
|
||||
|
||||
-- Check current connections
|
||||
select count(*) from pg_stat_activity; -- 487 connections!
|
||||
```
|
||||
|
||||
**Correct (connection pooling):**
|
||||
|
||||
```sql
|
||||
-- Use a pooler like PgBouncer between app and database
|
||||
-- Application connects to pooler, pooler reuses a small pool to Postgres
|
||||
|
||||
-- Configure pool_size based on: (CPU cores * 2) + spindle_count
|
||||
-- Example for 4 cores: pool_size = 10
|
||||
|
||||
-- Result: 500 concurrent users share 10 actual connections
|
||||
select count(*) from pg_stat_activity; -- 10 connections
|
||||
```
|
||||
|
||||
Pool modes:
|
||||
|
||||
- **Transaction mode**: connection returned after each transaction (best for most apps)
|
||||
- **Session mode**: connection held for entire session (needed for prepared statements, temp tables)
|
||||
|
||||
Reference: [Connection Pooling](https://supabase.com/docs/guides/database/connecting-to-postgres#connection-pooler)
|
||||
+46
@@ -0,0 +1,46 @@
|
||||
---
|
||||
title: Use Prepared Statements Correctly with Pooling
|
||||
impact: HIGH
|
||||
impactDescription: Avoid prepared statement conflicts in pooled environments
|
||||
tags: prepared-statements, connection-pooling, transaction-mode
|
||||
---
|
||||
|
||||
## Use Prepared Statements Correctly with Pooling
|
||||
|
||||
Prepared statements are tied to individual database connections. In transaction-mode pooling, connections are shared, causing conflicts.
|
||||
|
||||
**Incorrect (named prepared statements with transaction pooling):**
|
||||
|
||||
```sql
|
||||
-- Named prepared statement
|
||||
prepare get_user as select * from users where id = $1;
|
||||
|
||||
-- In transaction mode pooling, next request may get different connection
|
||||
execute get_user(123);
|
||||
-- ERROR: prepared statement "get_user" does not exist
|
||||
```
|
||||
|
||||
**Correct (use unnamed statements or session mode):**
|
||||
|
||||
```sql
|
||||
-- Option 1: Use unnamed prepared statements (most ORMs do this automatically)
|
||||
-- The query is prepared and executed in a single protocol message
|
||||
|
||||
-- Option 2: Deallocate after use in transaction mode
|
||||
prepare get_user as select * from users where id = $1;
|
||||
execute get_user(123);
|
||||
deallocate get_user;
|
||||
|
||||
-- Option 3: Use session mode pooling (port 5432 vs 6543)
|
||||
-- Connection is held for entire session, prepared statements persist
|
||||
```
|
||||
|
||||
Check your driver settings:
|
||||
|
||||
```sql
|
||||
-- Many drivers use prepared statements by default
|
||||
-- Node.js pg: { prepare: false } to disable
|
||||
-- JDBC: prepareThreshold=0 to disable
|
||||
```
|
||||
|
||||
Reference: [Prepared Statements with Pooling](https://supabase.com/docs/guides/database/connecting-to-postgres#connection-pool-modes)
|
||||
+54
@@ -0,0 +1,54 @@
|
||||
---
|
||||
title: Batch INSERT Statements for Bulk Data
|
||||
impact: MEDIUM
|
||||
impactDescription: 10-50x faster bulk inserts
|
||||
tags: batch, insert, bulk, performance, copy
|
||||
---
|
||||
|
||||
## Batch INSERT Statements for Bulk Data
|
||||
|
||||
Individual INSERT statements have high overhead. Batch multiple rows in single statements or use COPY.
|
||||
|
||||
**Incorrect (individual inserts):**
|
||||
|
||||
```sql
|
||||
-- Each insert is a separate transaction and round trip
|
||||
insert into events (user_id, action) values (1, 'click');
|
||||
insert into events (user_id, action) values (1, 'view');
|
||||
insert into events (user_id, action) values (2, 'click');
|
||||
-- ... 1000 more individual inserts
|
||||
|
||||
-- 1000 inserts = 1000 round trips = slow
|
||||
```
|
||||
|
||||
**Correct (batch insert):**
|
||||
|
||||
```sql
|
||||
-- Multiple rows in single statement
|
||||
insert into events (user_id, action) values
|
||||
(1, 'click'),
|
||||
(1, 'view'),
|
||||
(2, 'click'),
|
||||
-- ... up to ~1000 rows per batch
|
||||
(999, 'view');
|
||||
|
||||
-- One round trip for 1000 rows
|
||||
```
|
||||
|
||||
For large imports, use COPY:
|
||||
|
||||
```sql
|
||||
-- COPY is fastest for bulk loading
|
||||
copy events (user_id, action, created_at)
|
||||
from '/path/to/data.csv'
|
||||
with (format csv, header true);
|
||||
|
||||
-- Or from stdin in application
|
||||
copy events (user_id, action) from stdin with (format csv);
|
||||
1,click
|
||||
1,view
|
||||
2,click
|
||||
\.
|
||||
```
|
||||
|
||||
Reference: [COPY](https://www.postgresql.org/docs/current/sql-copy.html)
|
||||
+53
@@ -0,0 +1,53 @@
|
||||
---
|
||||
title: Eliminate N+1 Queries with Batch Loading
|
||||
impact: MEDIUM-HIGH
|
||||
impactDescription: 10-100x fewer database round trips
|
||||
tags: n-plus-one, batch, performance, queries
|
||||
---
|
||||
|
||||
## Eliminate N+1 Queries with Batch Loading
|
||||
|
||||
N+1 queries execute one query per item in a loop. Batch them into a single query using arrays or JOINs.
|
||||
|
||||
**Incorrect (N+1 queries):**
|
||||
|
||||
```sql
|
||||
-- First query: get all users
|
||||
select id from users where active = true; -- Returns 100 IDs
|
||||
|
||||
-- Then N queries, one per user
|
||||
select * from orders where user_id = 1;
|
||||
select * from orders where user_id = 2;
|
||||
select * from orders where user_id = 3;
|
||||
-- ... 97 more queries!
|
||||
|
||||
-- Total: 101 round trips to database
|
||||
```
|
||||
|
||||
**Correct (single batch query):**
|
||||
|
||||
```sql
|
||||
-- Collect IDs and query once with ANY
|
||||
select * from orders where user_id = any(array[1, 2, 3, ...]);
|
||||
|
||||
-- Or use JOIN instead of loop
|
||||
select u.id, u.name, o.*
|
||||
from users u
|
||||
left join orders o on o.user_id = u.id
|
||||
where u.active = true;
|
||||
|
||||
-- Total: 1 round trip
|
||||
```
|
||||
|
||||
Application pattern:
|
||||
|
||||
```sql
|
||||
-- Instead of looping in application code:
|
||||
-- for user in users: db.query("SELECT * FROM orders WHERE user_id = $1", user.id)
|
||||
|
||||
-- Pass array parameter:
|
||||
select * from orders where user_id = any($1::bigint[]);
|
||||
-- Application passes: [1, 2, 3, 4, 5, ...]
|
||||
```
|
||||
|
||||
Reference: [N+1 Query Problem](https://supabase.com/docs/guides/database/query-optimization)
|
||||
+50
@@ -0,0 +1,50 @@
|
||||
---
|
||||
title: Use Cursor-Based Pagination Instead of OFFSET
|
||||
impact: MEDIUM-HIGH
|
||||
impactDescription: Consistent O(1) performance regardless of page depth
|
||||
tags: pagination, cursor, keyset, offset, performance
|
||||
---
|
||||
|
||||
## Use Cursor-Based Pagination Instead of OFFSET
|
||||
|
||||
OFFSET-based pagination scans all skipped rows, getting slower on deeper pages. Cursor pagination is O(1).
|
||||
|
||||
**Incorrect (OFFSET pagination):**
|
||||
|
||||
```sql
|
||||
-- Page 1: scans 20 rows
|
||||
select * from products order by id limit 20 offset 0;
|
||||
|
||||
-- Page 100: scans 2000 rows to skip 1980
|
||||
select * from products order by id limit 20 offset 1980;
|
||||
|
||||
-- Page 10000: scans 200,000 rows!
|
||||
select * from products order by id limit 20 offset 199980;
|
||||
```
|
||||
|
||||
**Correct (cursor/keyset pagination):**
|
||||
|
||||
```sql
|
||||
-- Page 1: get first 20
|
||||
select * from products order by id limit 20;
|
||||
-- Application stores last_id = 20
|
||||
|
||||
-- Page 2: start after last ID
|
||||
select * from products where id > 20 order by id limit 20;
|
||||
-- Uses index, always fast regardless of page depth
|
||||
|
||||
-- Page 10000: same speed as page 1
|
||||
select * from products where id > 199980 order by id limit 20;
|
||||
```
|
||||
|
||||
For multi-column sorting:
|
||||
|
||||
```sql
|
||||
-- Cursor must include all sort columns
|
||||
select * from products
|
||||
where (created_at, id) > ('2024-01-15 10:00:00', 12345)
|
||||
order by created_at, id
|
||||
limit 20;
|
||||
```
|
||||
|
||||
Reference: [Pagination](https://supabase.com/docs/guides/database/pagination)
|
||||
+50
@@ -0,0 +1,50 @@
|
||||
---
|
||||
title: Use UPSERT for Insert-or-Update Operations
|
||||
impact: MEDIUM
|
||||
impactDescription: Atomic operation, eliminates race conditions
|
||||
tags: upsert, on-conflict, insert, update
|
||||
---
|
||||
|
||||
## Use UPSERT for Insert-or-Update Operations
|
||||
|
||||
Using separate SELECT-then-INSERT/UPDATE creates race conditions. Use INSERT ... ON CONFLICT for atomic upserts.
|
||||
|
||||
**Incorrect (check-then-insert race condition):**
|
||||
|
||||
```sql
|
||||
-- Race condition: two requests check simultaneously
|
||||
select * from settings where user_id = 123 and key = 'theme';
|
||||
-- Both find nothing
|
||||
|
||||
-- Both try to insert
|
||||
insert into settings (user_id, key, value) values (123, 'theme', 'dark');
|
||||
-- One succeeds, one fails with duplicate key error!
|
||||
```
|
||||
|
||||
**Correct (atomic UPSERT):**
|
||||
|
||||
```sql
|
||||
-- Single atomic operation
|
||||
insert into settings (user_id, key, value)
|
||||
values (123, 'theme', 'dark')
|
||||
on conflict (user_id, key)
|
||||
do update set value = excluded.value, updated_at = now();
|
||||
|
||||
-- Returns the inserted/updated row
|
||||
insert into settings (user_id, key, value)
|
||||
values (123, 'theme', 'dark')
|
||||
on conflict (user_id, key)
|
||||
do update set value = excluded.value
|
||||
returning *;
|
||||
```
|
||||
|
||||
Insert-or-ignore pattern:
|
||||
|
||||
```sql
|
||||
-- Insert only if not exists (no update)
|
||||
insert into page_views (page_id, user_id)
|
||||
values (1, 123)
|
||||
on conflict (page_id, user_id) do nothing;
|
||||
```
|
||||
|
||||
Reference: [INSERT ON CONFLICT](https://www.postgresql.org/docs/current/sql-insert.html#SQL-ON-CONFLICT)
|
||||
+56
@@ -0,0 +1,56 @@
|
||||
---
|
||||
title: Use Advisory Locks for Application-Level Locking
|
||||
impact: MEDIUM
|
||||
impactDescription: Efficient coordination without row-level lock overhead
|
||||
tags: advisory-locks, coordination, application-locks
|
||||
---
|
||||
|
||||
## Use Advisory Locks for Application-Level Locking
|
||||
|
||||
Advisory locks provide application-level coordination without requiring database rows to lock.
|
||||
|
||||
**Incorrect (creating rows just for locking):**
|
||||
|
||||
```sql
|
||||
-- Creating dummy rows to lock on
|
||||
create table resource_locks (
|
||||
resource_name text primary key
|
||||
);
|
||||
|
||||
insert into resource_locks values ('report_generator');
|
||||
|
||||
-- Lock by selecting the row
|
||||
select * from resource_locks where resource_name = 'report_generator' for update;
|
||||
```
|
||||
|
||||
**Correct (advisory locks):**
|
||||
|
||||
```sql
|
||||
-- Session-level advisory lock (released on disconnect or unlock)
|
||||
select pg_advisory_lock(hashtext('report_generator'));
|
||||
-- ... do exclusive work ...
|
||||
select pg_advisory_unlock(hashtext('report_generator'));
|
||||
|
||||
-- Transaction-level lock (released on commit/rollback)
|
||||
begin;
|
||||
select pg_advisory_xact_lock(hashtext('daily_report'));
|
||||
-- ... do work ...
|
||||
commit; -- Lock automatically released
|
||||
```
|
||||
|
||||
Try-lock for non-blocking operations:
|
||||
|
||||
```sql
|
||||
-- Returns immediately with true/false instead of waiting
|
||||
select pg_try_advisory_lock(hashtext('resource_name'));
|
||||
|
||||
-- Use in application
|
||||
if (acquired) {
|
||||
-- Do work
|
||||
select pg_advisory_unlock(hashtext('resource_name'));
|
||||
} else {
|
||||
-- Skip or retry later
|
||||
}
|
||||
```
|
||||
|
||||
Reference: [Advisory Locks](https://www.postgresql.org/docs/current/explicit-locking.html#ADVISORY-LOCKS)
|
||||
+68
@@ -0,0 +1,68 @@
|
||||
---
|
||||
title: Prevent Deadlocks with Consistent Lock Ordering
|
||||
impact: MEDIUM-HIGH
|
||||
impactDescription: Eliminate deadlock errors, improve reliability
|
||||
tags: deadlocks, locking, transactions, ordering
|
||||
---
|
||||
|
||||
## Prevent Deadlocks with Consistent Lock Ordering
|
||||
|
||||
Deadlocks occur when transactions lock resources in different orders. Always
|
||||
acquire locks in a consistent order.
|
||||
|
||||
**Incorrect (inconsistent lock ordering):**
|
||||
|
||||
```sql
|
||||
-- Transaction A -- Transaction B
|
||||
begin; begin;
|
||||
update accounts update accounts
|
||||
set balance = balance - 100 set balance = balance - 50
|
||||
where id = 1; where id = 2; -- B locks row 2
|
||||
|
||||
update accounts update accounts
|
||||
set balance = balance + 100 set balance = balance + 50
|
||||
where id = 2; -- A waits for B where id = 1; -- B waits for A
|
||||
|
||||
-- DEADLOCK! Both waiting for each other
|
||||
```
|
||||
|
||||
**Correct (lock rows in consistent order first):**
|
||||
|
||||
```sql
|
||||
-- Explicitly acquire locks in ID order before updating
|
||||
begin;
|
||||
select * from accounts where id in (1, 2) order by id for update;
|
||||
|
||||
-- Now perform updates in any order - locks already held
|
||||
update accounts set balance = balance - 100 where id = 1;
|
||||
update accounts set balance = balance + 100 where id = 2;
|
||||
commit;
|
||||
```
|
||||
|
||||
Alternative: use a single statement to update atomically:
|
||||
|
||||
```sql
|
||||
-- Single statement acquires all locks atomically
|
||||
begin;
|
||||
update accounts
|
||||
set balance = balance + case id
|
||||
when 1 then -100
|
||||
when 2 then 100
|
||||
end
|
||||
where id in (1, 2);
|
||||
commit;
|
||||
```
|
||||
|
||||
Detect deadlocks in logs:
|
||||
|
||||
```sql
|
||||
-- Check for recent deadlocks
|
||||
select * from pg_stat_database where deadlocks > 0;
|
||||
|
||||
-- Enable deadlock logging
|
||||
set log_lock_waits = on;
|
||||
set deadlock_timeout = '1s';
|
||||
```
|
||||
|
||||
Reference:
|
||||
[Deadlocks](https://www.postgresql.org/docs/current/explicit-locking.html#LOCKING-DEADLOCKS)
|
||||
+50
@@ -0,0 +1,50 @@
|
||||
---
|
||||
title: Keep Transactions Short to Reduce Lock Contention
|
||||
impact: MEDIUM-HIGH
|
||||
impactDescription: 3-5x throughput improvement, fewer deadlocks
|
||||
tags: transactions, locking, contention, performance
|
||||
---
|
||||
|
||||
## Keep Transactions Short to Reduce Lock Contention
|
||||
|
||||
Long-running transactions hold locks that block other queries. Keep transactions as short as possible.
|
||||
|
||||
**Incorrect (long transaction with external calls):**
|
||||
|
||||
```sql
|
||||
begin;
|
||||
select * from orders where id = 1 for update; -- Lock acquired
|
||||
|
||||
-- Application makes HTTP call to payment API (2-5 seconds)
|
||||
-- Other queries on this row are blocked!
|
||||
|
||||
update orders set status = 'paid' where id = 1;
|
||||
commit; -- Lock held for entire duration
|
||||
```
|
||||
|
||||
**Correct (minimal transaction scope):**
|
||||
|
||||
```sql
|
||||
-- Validate data and call APIs outside transaction
|
||||
-- Application: response = await paymentAPI.charge(...)
|
||||
|
||||
-- Only hold lock for the actual update
|
||||
begin;
|
||||
update orders
|
||||
set status = 'paid', payment_id = $1
|
||||
where id = $2 and status = 'pending'
|
||||
returning *;
|
||||
commit; -- Lock held for milliseconds
|
||||
```
|
||||
|
||||
Use `statement_timeout` to prevent runaway transactions:
|
||||
|
||||
```sql
|
||||
-- Abort queries running longer than 30 seconds
|
||||
set statement_timeout = '30s';
|
||||
|
||||
-- Or per-session
|
||||
set local statement_timeout = '5s';
|
||||
```
|
||||
|
||||
Reference: [Transaction Management](https://www.postgresql.org/docs/current/tutorial-transactions.html)
|
||||
+54
@@ -0,0 +1,54 @@
|
||||
---
|
||||
title: Use SKIP LOCKED for Non-Blocking Queue Processing
|
||||
impact: MEDIUM-HIGH
|
||||
impactDescription: 10x throughput for worker queues
|
||||
tags: skip-locked, queue, workers, concurrency
|
||||
---
|
||||
|
||||
## Use SKIP LOCKED for Non-Blocking Queue Processing
|
||||
|
||||
When multiple workers process a queue, SKIP LOCKED allows workers to process different rows without waiting.
|
||||
|
||||
**Incorrect (workers block each other):**
|
||||
|
||||
```sql
|
||||
-- Worker 1 and Worker 2 both try to get next job
|
||||
begin;
|
||||
select * from jobs where status = 'pending' order by created_at limit 1 for update;
|
||||
-- Worker 2 waits for Worker 1's lock to release!
|
||||
```
|
||||
|
||||
**Correct (SKIP LOCKED for parallel processing):**
|
||||
|
||||
```sql
|
||||
-- Each worker skips locked rows and gets the next available
|
||||
begin;
|
||||
select * from jobs
|
||||
where status = 'pending'
|
||||
order by created_at
|
||||
limit 1
|
||||
for update skip locked;
|
||||
|
||||
-- Worker 1 gets job 1, Worker 2 gets job 2 (no waiting)
|
||||
|
||||
update jobs set status = 'processing' where id = $1;
|
||||
commit;
|
||||
```
|
||||
|
||||
Complete queue pattern:
|
||||
|
||||
```sql
|
||||
-- Atomic claim-and-update in one statement
|
||||
update jobs
|
||||
set status = 'processing', worker_id = $1, started_at = now()
|
||||
where id = (
|
||||
select id from jobs
|
||||
where status = 'pending'
|
||||
order by created_at
|
||||
limit 1
|
||||
for update skip locked
|
||||
)
|
||||
returning *;
|
||||
```
|
||||
|
||||
Reference: [SELECT FOR UPDATE SKIP LOCKED](https://www.postgresql.org/docs/current/sql-select.html#SQL-FOR-UPDATE-SHARE)
|
||||
+45
@@ -0,0 +1,45 @@
|
||||
---
|
||||
title: Use EXPLAIN ANALYZE to Diagnose Slow Queries
|
||||
impact: LOW-MEDIUM
|
||||
impactDescription: Identify exact bottlenecks in query execution
|
||||
tags: explain, analyze, diagnostics, query-plan
|
||||
---
|
||||
|
||||
## Use EXPLAIN ANALYZE to Diagnose Slow Queries
|
||||
|
||||
EXPLAIN ANALYZE executes the query and shows actual timings, revealing the true performance bottlenecks.
|
||||
|
||||
**Incorrect (guessing at performance issues):**
|
||||
|
||||
```sql
|
||||
-- Query is slow, but why?
|
||||
select * from orders where customer_id = 123 and status = 'pending';
|
||||
-- "It must be missing an index" - but which one?
|
||||
```
|
||||
|
||||
**Correct (use EXPLAIN ANALYZE):**
|
||||
|
||||
```sql
|
||||
explain (analyze, buffers, format text)
|
||||
select * from orders where customer_id = 123 and status = 'pending';
|
||||
|
||||
-- Output reveals the issue:
|
||||
-- Seq Scan on orders (cost=0.00..25000.00 rows=50 width=100) (actual time=0.015..450.123 rows=50 loops=1)
|
||||
-- Filter: ((customer_id = 123) AND (status = 'pending'::text))
|
||||
-- Rows Removed by Filter: 999950
|
||||
-- Buffers: shared hit=5000 read=15000
|
||||
-- Planning Time: 0.150 ms
|
||||
-- Execution Time: 450.500 ms
|
||||
```
|
||||
|
||||
Key things to look for:
|
||||
|
||||
```sql
|
||||
-- Seq Scan on large tables = missing index
|
||||
-- Rows Removed by Filter = poor selectivity or missing index
|
||||
-- Buffers: read >> hit = data not cached, needs more memory
|
||||
-- Nested Loop with high loops = consider different join strategy
|
||||
-- Sort Method: external merge = work_mem too low
|
||||
```
|
||||
|
||||
Reference: [EXPLAIN](https://supabase.com/docs/guides/database/inspect)
|
||||
+55
@@ -0,0 +1,55 @@
|
||||
---
|
||||
title: Enable pg_stat_statements for Query Analysis
|
||||
impact: LOW-MEDIUM
|
||||
impactDescription: Identify top resource-consuming queries
|
||||
tags: pg-stat-statements, monitoring, statistics, performance
|
||||
---
|
||||
|
||||
## Enable pg_stat_statements for Query Analysis
|
||||
|
||||
pg_stat_statements tracks execution statistics for all queries, helping identify slow and frequent queries.
|
||||
|
||||
**Incorrect (no visibility into query patterns):**
|
||||
|
||||
```sql
|
||||
-- Database is slow, but which queries are the problem?
|
||||
-- No way to know without pg_stat_statements
|
||||
```
|
||||
|
||||
**Correct (enable and query pg_stat_statements):**
|
||||
|
||||
```sql
|
||||
-- Enable the extension
|
||||
create extension if not exists pg_stat_statements;
|
||||
|
||||
-- Find slowest queries by total time
|
||||
select
|
||||
calls,
|
||||
round(total_exec_time::numeric, 2) as total_time_ms,
|
||||
round(mean_exec_time::numeric, 2) as mean_time_ms,
|
||||
query
|
||||
from pg_stat_statements
|
||||
order by total_exec_time desc
|
||||
limit 10;
|
||||
|
||||
-- Find most frequent queries
|
||||
select calls, query
|
||||
from pg_stat_statements
|
||||
order by calls desc
|
||||
limit 10;
|
||||
|
||||
-- Reset statistics after optimization
|
||||
select pg_stat_statements_reset();
|
||||
```
|
||||
|
||||
Key metrics to monitor:
|
||||
|
||||
```sql
|
||||
-- Queries with high mean time (candidates for optimization)
|
||||
select query, mean_exec_time, calls
|
||||
from pg_stat_statements
|
||||
where mean_exec_time > 100 -- > 100ms average
|
||||
order by mean_exec_time desc;
|
||||
```
|
||||
|
||||
Reference: [pg_stat_statements](https://supabase.com/docs/guides/database/extensions/pg_stat_statements)
|
||||
+55
@@ -0,0 +1,55 @@
|
||||
---
|
||||
title: Maintain Table Statistics with VACUUM and ANALYZE
|
||||
impact: MEDIUM
|
||||
impactDescription: 2-10x better query plans with accurate statistics
|
||||
tags: vacuum, analyze, statistics, maintenance, autovacuum
|
||||
---
|
||||
|
||||
## Maintain Table Statistics with VACUUM and ANALYZE
|
||||
|
||||
Outdated statistics cause the query planner to make poor decisions. VACUUM reclaims space, ANALYZE updates statistics.
|
||||
|
||||
**Incorrect (stale statistics):**
|
||||
|
||||
```sql
|
||||
-- Table has 1M rows but stats say 1000
|
||||
-- Query planner chooses wrong strategy
|
||||
explain select * from orders where status = 'pending';
|
||||
-- Shows: Seq Scan (because stats show small table)
|
||||
-- Actually: Index Scan would be much faster
|
||||
```
|
||||
|
||||
**Correct (maintain fresh statistics):**
|
||||
|
||||
```sql
|
||||
-- Manually analyze after large data changes
|
||||
analyze orders;
|
||||
|
||||
-- Analyze specific columns used in WHERE clauses
|
||||
analyze orders (status, created_at);
|
||||
|
||||
-- Check when tables were last analyzed
|
||||
select
|
||||
relname,
|
||||
last_vacuum,
|
||||
last_autovacuum,
|
||||
last_analyze,
|
||||
last_autoanalyze
|
||||
from pg_stat_user_tables
|
||||
order by last_analyze nulls first;
|
||||
```
|
||||
|
||||
Autovacuum tuning for busy tables:
|
||||
|
||||
```sql
|
||||
-- Increase frequency for high-churn tables
|
||||
alter table orders set (
|
||||
autovacuum_vacuum_scale_factor = 0.05, -- Vacuum at 5% dead tuples (default 20%)
|
||||
autovacuum_analyze_scale_factor = 0.02 -- Analyze at 2% changes (default 10%)
|
||||
);
|
||||
|
||||
-- Check autovacuum status
|
||||
select * from pg_stat_progress_vacuum;
|
||||
```
|
||||
|
||||
Reference: [VACUUM](https://supabase.com/docs/guides/database/database-size#vacuum-operations)
|
||||
+44
@@ -0,0 +1,44 @@
|
||||
---
|
||||
title: Create Composite Indexes for Multi-Column Queries
|
||||
impact: HIGH
|
||||
impactDescription: 5-10x faster multi-column queries
|
||||
tags: indexes, composite-index, multi-column, query-optimization
|
||||
---
|
||||
|
||||
## Create Composite Indexes for Multi-Column Queries
|
||||
|
||||
When queries filter on multiple columns, a composite index is more efficient than separate single-column indexes.
|
||||
|
||||
**Incorrect (separate indexes require bitmap scan):**
|
||||
|
||||
```sql
|
||||
-- Two separate indexes
|
||||
create index orders_status_idx on orders (status);
|
||||
create index orders_created_idx on orders (created_at);
|
||||
|
||||
-- Query must combine both indexes (slower)
|
||||
select * from orders where status = 'pending' and created_at > '2024-01-01';
|
||||
```
|
||||
|
||||
**Correct (composite index):**
|
||||
|
||||
```sql
|
||||
-- Single composite index (leftmost column first for equality checks)
|
||||
create index orders_status_created_idx on orders (status, created_at);
|
||||
|
||||
-- Query uses one efficient index scan
|
||||
select * from orders where status = 'pending' and created_at > '2024-01-01';
|
||||
```
|
||||
|
||||
**Column order matters** - place equality columns first, range columns last:
|
||||
|
||||
```sql
|
||||
-- Good: status (=) before created_at (>)
|
||||
create index idx on orders (status, created_at);
|
||||
|
||||
-- Works for: WHERE status = 'pending'
|
||||
-- Works for: WHERE status = 'pending' AND created_at > '2024-01-01'
|
||||
-- Does NOT work for: WHERE created_at > '2024-01-01' (leftmost prefix rule)
|
||||
```
|
||||
|
||||
Reference: [Multicolumn Indexes](https://www.postgresql.org/docs/current/indexes-multicolumn.html)
|
||||
+40
@@ -0,0 +1,40 @@
|
||||
---
|
||||
title: Use Covering Indexes to Avoid Table Lookups
|
||||
impact: MEDIUM-HIGH
|
||||
impactDescription: 2-5x faster queries by eliminating heap fetches
|
||||
tags: indexes, covering-index, include, index-only-scan
|
||||
---
|
||||
|
||||
## Use Covering Indexes to Avoid Table Lookups
|
||||
|
||||
Covering indexes include all columns needed by a query, enabling index-only scans that skip the table entirely.
|
||||
|
||||
**Incorrect (index scan + heap fetch):**
|
||||
|
||||
```sql
|
||||
create index users_email_idx on users (email);
|
||||
|
||||
-- Must fetch name and created_at from table heap
|
||||
select email, name, created_at from users where email = 'user@example.com';
|
||||
```
|
||||
|
||||
**Correct (index-only scan with INCLUDE):**
|
||||
|
||||
```sql
|
||||
-- Include non-searchable columns in the index
|
||||
create index users_email_idx on users (email) include (name, created_at);
|
||||
|
||||
-- All columns served from index, no table access needed
|
||||
select email, name, created_at from users where email = 'user@example.com';
|
||||
```
|
||||
|
||||
Use INCLUDE for columns you SELECT but don't filter on:
|
||||
|
||||
```sql
|
||||
-- Searching by status, but also need customer_id and total
|
||||
create index orders_status_idx on orders (status) include (customer_id, total);
|
||||
|
||||
select status, customer_id, total from orders where status = 'shipped';
|
||||
```
|
||||
|
||||
Reference: [Index-Only Scans](https://www.postgresql.org/docs/current/indexes-index-only-scans.html)
|
||||
+45
@@ -0,0 +1,45 @@
|
||||
---
|
||||
title: Choose the Right Index Type for Your Data
|
||||
impact: HIGH
|
||||
impactDescription: 10-100x improvement with correct index type
|
||||
tags: indexes, btree, gin, brin, hash, index-types
|
||||
---
|
||||
|
||||
## Choose the Right Index Type for Your Data
|
||||
|
||||
Different index types excel at different query patterns. The default B-tree isn't always optimal.
|
||||
|
||||
**Incorrect (B-tree for JSONB containment):**
|
||||
|
||||
```sql
|
||||
-- B-tree cannot optimize containment operators
|
||||
create index products_attrs_idx on products (attributes);
|
||||
select * from products where attributes @> '{"color": "red"}';
|
||||
-- Full table scan - B-tree doesn't support @> operator
|
||||
```
|
||||
|
||||
**Correct (GIN for JSONB):**
|
||||
|
||||
```sql
|
||||
-- GIN supports @>, ?, ?&, ?| operators
|
||||
create index products_attrs_idx on products using gin (attributes);
|
||||
select * from products where attributes @> '{"color": "red"}';
|
||||
```
|
||||
|
||||
Index type guide:
|
||||
|
||||
```sql
|
||||
-- B-tree (default): =, <, >, BETWEEN, IN, IS NULL
|
||||
create index users_created_idx on users (created_at);
|
||||
|
||||
-- GIN: arrays, JSONB, full-text search
|
||||
create index posts_tags_idx on posts using gin (tags);
|
||||
|
||||
-- BRIN: large time-series tables (10-100x smaller)
|
||||
create index events_time_idx on events using brin (created_at);
|
||||
|
||||
-- Hash: equality-only (slightly faster than B-tree for =)
|
||||
create index sessions_token_idx on sessions using hash (token);
|
||||
```
|
||||
|
||||
Reference: [Index Types](https://www.postgresql.org/docs/current/indexes-types.html)
|
||||
+43
@@ -0,0 +1,43 @@
|
||||
---
|
||||
title: Add Indexes on WHERE and JOIN Columns
|
||||
impact: CRITICAL
|
||||
impactDescription: 100-1000x faster queries on large tables
|
||||
tags: indexes, performance, sequential-scan, query-optimization
|
||||
---
|
||||
|
||||
## Add Indexes on WHERE and JOIN Columns
|
||||
|
||||
Queries filtering or joining on unindexed columns cause full table scans, which become exponentially slower as tables grow.
|
||||
|
||||
**Incorrect (sequential scan on large table):**
|
||||
|
||||
```sql
|
||||
-- No index on customer_id causes full table scan
|
||||
select * from orders where customer_id = 123;
|
||||
|
||||
-- EXPLAIN shows: Seq Scan on orders (cost=0.00..25000.00 rows=100 width=85)
|
||||
```
|
||||
|
||||
**Correct (index scan):**
|
||||
|
||||
```sql
|
||||
-- Create index on frequently filtered column
|
||||
create index orders_customer_id_idx on orders (customer_id);
|
||||
|
||||
select * from orders where customer_id = 123;
|
||||
|
||||
-- EXPLAIN shows: Index Scan using orders_customer_id_idx (cost=0.42..8.44 rows=100 width=85)
|
||||
```
|
||||
|
||||
For JOIN columns, always index the foreign key side:
|
||||
|
||||
```sql
|
||||
-- Index the referencing column
|
||||
create index orders_customer_id_idx on orders (customer_id);
|
||||
|
||||
select c.name, o.total
|
||||
from customers c
|
||||
join orders o on o.customer_id = c.id;
|
||||
```
|
||||
|
||||
Reference: [Query Optimization](https://supabase.com/docs/guides/database/query-optimization)
|
||||
+45
@@ -0,0 +1,45 @@
|
||||
---
|
||||
title: Use Partial Indexes for Filtered Queries
|
||||
impact: HIGH
|
||||
impactDescription: 5-20x smaller indexes, faster writes and queries
|
||||
tags: indexes, partial-index, query-optimization, storage
|
||||
---
|
||||
|
||||
## Use Partial Indexes for Filtered Queries
|
||||
|
||||
Partial indexes only include rows matching a WHERE condition, making them smaller and faster when queries consistently filter on the same condition.
|
||||
|
||||
**Incorrect (full index includes irrelevant rows):**
|
||||
|
||||
```sql
|
||||
-- Index includes all rows, even soft-deleted ones
|
||||
create index users_email_idx on users (email);
|
||||
|
||||
-- Query always filters active users
|
||||
select * from users where email = 'user@example.com' and deleted_at is null;
|
||||
```
|
||||
|
||||
**Correct (partial index matches query filter):**
|
||||
|
||||
```sql
|
||||
-- Index only includes active users
|
||||
create index users_active_email_idx on users (email)
|
||||
where deleted_at is null;
|
||||
|
||||
-- Query uses the smaller, faster index
|
||||
select * from users where email = 'user@example.com' and deleted_at is null;
|
||||
```
|
||||
|
||||
Common use cases for partial indexes:
|
||||
|
||||
```sql
|
||||
-- Only pending orders (status rarely changes once completed)
|
||||
create index orders_pending_idx on orders (created_at)
|
||||
where status = 'pending';
|
||||
|
||||
-- Only non-null values
|
||||
create index products_sku_idx on products (sku)
|
||||
where sku is not null;
|
||||
```
|
||||
|
||||
Reference: [Partial Indexes](https://www.postgresql.org/docs/current/indexes-partial.html)
|
||||
+46
@@ -0,0 +1,46 @@
|
||||
---
|
||||
title: Choose Appropriate Data Types
|
||||
impact: HIGH
|
||||
impactDescription: 50% storage reduction, faster comparisons
|
||||
tags: data-types, schema, storage, performance
|
||||
---
|
||||
|
||||
## Choose Appropriate Data Types
|
||||
|
||||
Using the right data types reduces storage, improves query performance, and prevents bugs.
|
||||
|
||||
**Incorrect (wrong data types):**
|
||||
|
||||
```sql
|
||||
create table users (
|
||||
id int, -- Will overflow at 2.1 billion
|
||||
email varchar(255), -- Unnecessary length limit
|
||||
created_at timestamp, -- Missing timezone info
|
||||
is_active varchar(5), -- String for boolean
|
||||
price varchar(20) -- String for numeric
|
||||
);
|
||||
```
|
||||
|
||||
**Correct (appropriate data types):**
|
||||
|
||||
```sql
|
||||
create table users (
|
||||
id bigint generated always as identity primary key, -- 9 quintillion max
|
||||
email text, -- No artificial limit, same performance as varchar
|
||||
created_at timestamptz, -- Always store timezone-aware timestamps
|
||||
is_active boolean default true, -- 1 byte vs variable string length
|
||||
price numeric(10,2) -- Exact decimal arithmetic
|
||||
);
|
||||
```
|
||||
|
||||
Key guidelines:
|
||||
|
||||
```sql
|
||||
-- IDs: use bigint, not int (future-proofing)
|
||||
-- Strings: use text, not varchar(n) unless constraint needed
|
||||
-- Time: use timestamptz, not timestamp
|
||||
-- Money: use numeric, not float (precision matters)
|
||||
-- Enums: use text with check constraint or create enum type
|
||||
```
|
||||
|
||||
Reference: [Data Types](https://www.postgresql.org/docs/current/datatype.html)
|
||||
+59
@@ -0,0 +1,59 @@
|
||||
---
|
||||
title: Index Foreign Key Columns
|
||||
impact: HIGH
|
||||
impactDescription: 10-100x faster JOINs and CASCADE operations
|
||||
tags: foreign-key, indexes, joins, schema
|
||||
---
|
||||
|
||||
## Index Foreign Key Columns
|
||||
|
||||
Postgres does not automatically index foreign key columns. Missing indexes cause slow JOINs and CASCADE operations.
|
||||
|
||||
**Incorrect (unindexed foreign key):**
|
||||
|
||||
```sql
|
||||
create table orders (
|
||||
id bigint generated always as identity primary key,
|
||||
customer_id bigint references customers(id) on delete cascade,
|
||||
total numeric(10,2)
|
||||
);
|
||||
|
||||
-- No index on customer_id!
|
||||
-- JOINs and ON DELETE CASCADE both require full table scan
|
||||
select * from orders where customer_id = 123; -- Seq Scan
|
||||
delete from customers where id = 123; -- Locks table, scans all orders
|
||||
```
|
||||
|
||||
**Correct (indexed foreign key):**
|
||||
|
||||
```sql
|
||||
create table orders (
|
||||
id bigint generated always as identity primary key,
|
||||
customer_id bigint references customers(id) on delete cascade,
|
||||
total numeric(10,2)
|
||||
);
|
||||
|
||||
-- Always index the FK column
|
||||
create index orders_customer_id_idx on orders (customer_id);
|
||||
|
||||
-- Now JOINs and cascades are fast
|
||||
select * from orders where customer_id = 123; -- Index Scan
|
||||
delete from customers where id = 123; -- Uses index, fast cascade
|
||||
```
|
||||
|
||||
Find missing FK indexes:
|
||||
|
||||
```sql
|
||||
select
|
||||
conrelid::regclass as table_name,
|
||||
a.attname as fk_column
|
||||
from pg_constraint c
|
||||
join pg_attribute a on a.attrelid = c.conrelid and a.attnum = any(c.conkey)
|
||||
where c.contype = 'f'
|
||||
and not exists (
|
||||
select 1 from pg_index i
|
||||
where i.indrelid = c.conrelid and a.attnum = any(i.indkey)
|
||||
);
|
||||
```
|
||||
|
||||
Reference: [Foreign Keys](https://www.postgresql.org/docs/current/ddl-constraints.html#DDL-CONSTRAINTS-FK)
|
||||
+55
@@ -0,0 +1,55 @@
|
||||
---
|
||||
title: Use Lowercase Identifiers for Compatibility
|
||||
impact: MEDIUM
|
||||
impactDescription: Avoid case-sensitivity bugs with tools, ORMs, and AI assistants
|
||||
tags: naming, identifiers, case-sensitivity, schema, conventions
|
||||
---
|
||||
|
||||
## Use Lowercase Identifiers for Compatibility
|
||||
|
||||
PostgreSQL folds unquoted identifiers to lowercase. Quoted mixed-case identifiers require quotes forever and cause issues with tools, ORMs, and AI assistants that may not recognize them.
|
||||
|
||||
**Incorrect (mixed-case identifiers):**
|
||||
|
||||
```sql
|
||||
-- Quoted identifiers preserve case but require quotes everywhere
|
||||
CREATE TABLE "Users" (
|
||||
"userId" bigint PRIMARY KEY,
|
||||
"firstName" text,
|
||||
"lastName" text
|
||||
);
|
||||
|
||||
-- Must always quote or queries fail
|
||||
SELECT "firstName" FROM "Users" WHERE "userId" = 1;
|
||||
|
||||
-- This fails - Users becomes users without quotes
|
||||
SELECT firstName FROM Users;
|
||||
-- ERROR: relation "users" does not exist
|
||||
```
|
||||
|
||||
**Correct (lowercase snake_case):**
|
||||
|
||||
```sql
|
||||
-- Unquoted lowercase identifiers are portable and tool-friendly
|
||||
CREATE TABLE users (
|
||||
user_id bigint PRIMARY KEY,
|
||||
first_name text,
|
||||
last_name text
|
||||
);
|
||||
|
||||
-- Works without quotes, recognized by all tools
|
||||
SELECT first_name FROM users WHERE user_id = 1;
|
||||
```
|
||||
|
||||
Common sources of mixed-case identifiers:
|
||||
|
||||
```sql
|
||||
-- ORMs often generate quoted camelCase - configure them to use snake_case
|
||||
-- Migrations from other databases may preserve original casing
|
||||
-- Some GUI tools quote identifiers by default - disable this
|
||||
|
||||
-- If stuck with mixed-case, create views as a compatibility layer
|
||||
CREATE VIEW users AS SELECT "userId" AS user_id, "firstName" AS first_name FROM "Users";
|
||||
```
|
||||
|
||||
Reference: [Identifiers and Key Words](https://www.postgresql.org/docs/current/sql-syntax-lexical.html#SQL-SYNTAX-IDENTIFIERS)
|
||||
+55
@@ -0,0 +1,55 @@
|
||||
---
|
||||
title: Partition Large Tables for Better Performance
|
||||
impact: MEDIUM-HIGH
|
||||
impactDescription: 5-20x faster queries and maintenance on large tables
|
||||
tags: partitioning, large-tables, time-series, performance
|
||||
---
|
||||
|
||||
## Partition Large Tables for Better Performance
|
||||
|
||||
Partitioning splits a large table into smaller pieces, improving query performance and maintenance operations.
|
||||
|
||||
**Incorrect (single large table):**
|
||||
|
||||
```sql
|
||||
create table events (
|
||||
id bigint generated always as identity,
|
||||
created_at timestamptz,
|
||||
data jsonb
|
||||
);
|
||||
|
||||
-- 500M rows, queries scan everything
|
||||
select * from events where created_at > '2024-01-01'; -- Slow
|
||||
vacuum events; -- Takes hours, locks table
|
||||
```
|
||||
|
||||
**Correct (partitioned by time range):**
|
||||
|
||||
```sql
|
||||
create table events (
|
||||
id bigint generated always as identity,
|
||||
created_at timestamptz not null,
|
||||
data jsonb
|
||||
) partition by range (created_at);
|
||||
|
||||
-- Create partitions for each month
|
||||
create table events_2024_01 partition of events
|
||||
for values from ('2024-01-01') to ('2024-02-01');
|
||||
|
||||
create table events_2024_02 partition of events
|
||||
for values from ('2024-02-01') to ('2024-03-01');
|
||||
|
||||
-- Queries only scan relevant partitions
|
||||
select * from events where created_at > '2024-01-15'; -- Only scans events_2024_01+
|
||||
|
||||
-- Drop old data instantly
|
||||
drop table events_2023_01; -- Instant vs DELETE taking hours
|
||||
```
|
||||
|
||||
When to partition:
|
||||
|
||||
- Tables > 100M rows
|
||||
- Time-series data with date-based queries
|
||||
- Need to efficiently drop old data
|
||||
|
||||
Reference: [Table Partitioning](https://www.postgresql.org/docs/current/ddl-partitioning.html)
|
||||
+61
@@ -0,0 +1,61 @@
|
||||
---
|
||||
title: Select Optimal Primary Key Strategy
|
||||
impact: HIGH
|
||||
impactDescription: Better index locality, reduced fragmentation
|
||||
tags: primary-key, identity, uuid, serial, schema
|
||||
---
|
||||
|
||||
## Select Optimal Primary Key Strategy
|
||||
|
||||
Primary key choice affects insert performance, index size, and replication
|
||||
efficiency.
|
||||
|
||||
**Incorrect (problematic PK choices):**
|
||||
|
||||
```sql
|
||||
-- identity is the SQL-standard approach
|
||||
create table users (
|
||||
id serial primary key -- Works, but IDENTITY is recommended
|
||||
);
|
||||
|
||||
-- Random UUIDs (v4) cause index fragmentation
|
||||
create table orders (
|
||||
id uuid default gen_random_uuid() primary key -- UUIDv4 = random = scattered inserts
|
||||
);
|
||||
```
|
||||
|
||||
**Correct (optimal PK strategies):**
|
||||
|
||||
```sql
|
||||
-- Use IDENTITY for sequential IDs (SQL-standard, best for most cases)
|
||||
create table users (
|
||||
id bigint generated always as identity primary key
|
||||
);
|
||||
|
||||
-- For distributed systems needing UUIDs, use UUIDv7 (time-ordered)
|
||||
-- Requires pg_uuidv7 extension: create extension pg_uuidv7;
|
||||
create table orders (
|
||||
id uuid default uuid_generate_v7() primary key -- Time-ordered, no fragmentation
|
||||
);
|
||||
|
||||
-- Alternative: time-prefixed IDs for sortable, distributed IDs (no extension needed)
|
||||
create table events (
|
||||
id text default concat(
|
||||
to_char(now() at time zone 'utc', 'YYYYMMDDHH24MISSMS'),
|
||||
gen_random_uuid()::text
|
||||
) primary key
|
||||
);
|
||||
```
|
||||
|
||||
Guidelines:
|
||||
|
||||
- Single database: `bigint identity` (sequential, 8 bytes, SQL-standard)
|
||||
- Distributed/exposed IDs: UUIDv7 (requires pg_uuidv7) or ULID (time-ordered, no
|
||||
fragmentation)
|
||||
- `serial` works but `identity` is SQL-standard and preferred for new
|
||||
applications
|
||||
- Avoid random UUIDs (v4) as primary keys on large tables (causes index
|
||||
fragmentation)
|
||||
|
||||
Reference:
|
||||
[Identity Columns](https://www.postgresql.org/docs/current/sql-createtable.html#SQL-CREATETABLE-PARMS-GENERATED-IDENTITY)
|
||||
+54
@@ -0,0 +1,54 @@
|
||||
---
|
||||
title: Apply Principle of Least Privilege
|
||||
impact: MEDIUM
|
||||
impactDescription: Reduced attack surface, better audit trail
|
||||
tags: privileges, security, roles, permissions
|
||||
---
|
||||
|
||||
## Apply Principle of Least Privilege
|
||||
|
||||
Grant only the minimum permissions required. Never use superuser for application queries.
|
||||
|
||||
**Incorrect (overly broad permissions):**
|
||||
|
||||
```sql
|
||||
-- Application uses superuser connection
|
||||
-- Or grants ALL to application role
|
||||
grant all privileges on all tables in schema public to app_user;
|
||||
grant all privileges on all sequences in schema public to app_user;
|
||||
|
||||
-- Any SQL injection becomes catastrophic
|
||||
-- drop table users; cascades to everything
|
||||
```
|
||||
|
||||
**Correct (minimal, specific grants):**
|
||||
|
||||
```sql
|
||||
-- Create role with no default privileges
|
||||
create role app_readonly nologin;
|
||||
|
||||
-- Grant only SELECT on specific tables
|
||||
grant usage on schema public to app_readonly;
|
||||
grant select on public.products, public.categories to app_readonly;
|
||||
|
||||
-- Create role for writes with limited scope
|
||||
create role app_writer nologin;
|
||||
grant usage on schema public to app_writer;
|
||||
grant select, insert, update on public.orders to app_writer;
|
||||
grant usage on sequence orders_id_seq to app_writer;
|
||||
-- No DELETE permission
|
||||
|
||||
-- Login role inherits from these
|
||||
create role app_user login password 'xxx';
|
||||
grant app_writer to app_user;
|
||||
```
|
||||
|
||||
Revoke public defaults:
|
||||
|
||||
```sql
|
||||
-- Revoke default public access
|
||||
revoke all on schema public from public;
|
||||
revoke all on all tables in schema public from public;
|
||||
```
|
||||
|
||||
Reference: [Roles and Privileges](https://supabase.com/blog/postgres-roles-and-privileges)
|
||||
+50
@@ -0,0 +1,50 @@
|
||||
---
|
||||
title: Enable Row Level Security for Multi-Tenant Data
|
||||
impact: CRITICAL
|
||||
impactDescription: Database-enforced tenant isolation, prevent data leaks
|
||||
tags: rls, row-level-security, multi-tenant, security
|
||||
---
|
||||
|
||||
## Enable Row Level Security for Multi-Tenant Data
|
||||
|
||||
Row Level Security (RLS) enforces data access at the database level, ensuring users only see their own data.
|
||||
|
||||
**Incorrect (application-level filtering only):**
|
||||
|
||||
```sql
|
||||
-- Relying only on application to filter
|
||||
select * from orders where user_id = $current_user_id;
|
||||
|
||||
-- Bug or bypass means all data is exposed!
|
||||
select * from orders; -- Returns ALL orders
|
||||
```
|
||||
|
||||
**Correct (database-enforced RLS):**
|
||||
|
||||
```sql
|
||||
-- Enable RLS on the table
|
||||
alter table orders enable row level security;
|
||||
|
||||
-- Create policy for users to see only their orders
|
||||
create policy orders_user_policy on orders
|
||||
for all
|
||||
using (user_id = current_setting('app.current_user_id')::bigint);
|
||||
|
||||
-- Force RLS even for table owners
|
||||
alter table orders force row level security;
|
||||
|
||||
-- Set user context and query
|
||||
set app.current_user_id = '123';
|
||||
select * from orders; -- Only returns orders for user 123
|
||||
```
|
||||
|
||||
Policy for authenticated role:
|
||||
|
||||
```sql
|
||||
create policy orders_user_policy on orders
|
||||
for all
|
||||
to authenticated
|
||||
using (user_id = auth.uid());
|
||||
```
|
||||
|
||||
Reference: [Row Level Security](https://supabase.com/docs/guides/database/postgres/row-level-security)
|
||||
+57
@@ -0,0 +1,57 @@
|
||||
---
|
||||
title: Optimize RLS Policies for Performance
|
||||
impact: HIGH
|
||||
impactDescription: 5-10x faster RLS queries with proper patterns
|
||||
tags: rls, performance, security, optimization
|
||||
---
|
||||
|
||||
## Optimize RLS Policies for Performance
|
||||
|
||||
Poorly written RLS policies can cause severe performance issues. Use subqueries and indexes strategically.
|
||||
|
||||
**Incorrect (function called for every row):**
|
||||
|
||||
```sql
|
||||
create policy orders_policy on orders
|
||||
using (auth.uid() = user_id); -- auth.uid() called per row!
|
||||
|
||||
-- With 1M rows, auth.uid() is called 1M times
|
||||
```
|
||||
|
||||
**Correct (wrap functions in SELECT):**
|
||||
|
||||
```sql
|
||||
create policy orders_policy on orders
|
||||
using ((select auth.uid()) = user_id); -- Called once, cached
|
||||
|
||||
-- 100x+ faster on large tables
|
||||
```
|
||||
|
||||
Use security definer functions for complex checks:
|
||||
|
||||
```sql
|
||||
-- Create helper function (runs as definer, bypasses RLS)
|
||||
create or replace function is_team_member(team_id bigint)
|
||||
returns boolean
|
||||
language sql
|
||||
security definer
|
||||
set search_path = ''
|
||||
as $$
|
||||
select exists (
|
||||
select 1 from public.team_members
|
||||
where team_id = $1 and user_id = (select auth.uid())
|
||||
);
|
||||
$$;
|
||||
|
||||
-- Use in policy (indexed lookup, not per-row check)
|
||||
create policy team_orders_policy on orders
|
||||
using ((select is_team_member(team_id)));
|
||||
```
|
||||
|
||||
Always add indexes on columns used in RLS policies:
|
||||
|
||||
```sql
|
||||
create index orders_user_id_idx on orders (user_id);
|
||||
```
|
||||
|
||||
Reference: [RLS Performance](https://supabase.com/docs/guides/database/postgres/row-level-security#rls-performance-recommendations)
|
||||
+341
@@ -0,0 +1,341 @@
|
||||
---
|
||||
name: rag-engineer
|
||||
description: Expert in building Retrieval-Augmented Generation systems. Masters
|
||||
embedding models, vector databases, chunking strategies, and retrieval
|
||||
optimization for LLM applications.
|
||||
risk: unknown
|
||||
source: vibeship-spawner-skills (Apache 2.0)
|
||||
date_added: 2026-02-27
|
||||
---
|
||||
|
||||
# RAG Engineer
|
||||
|
||||
Expert in building Retrieval-Augmented Generation systems. Masters embedding models,
|
||||
vector databases, chunking strategies, and retrieval optimization for LLM applications.
|
||||
|
||||
**Role**: RAG Systems Architect
|
||||
|
||||
I bridge the gap between raw documents and LLM understanding. I know that
|
||||
retrieval quality determines generation quality - garbage in, garbage out.
|
||||
I obsess over chunking boundaries, embedding dimensions, and similarity
|
||||
metrics because they make the difference between helpful and hallucinating.
|
||||
|
||||
### Expertise
|
||||
|
||||
- Embedding model selection and fine-tuning
|
||||
- Vector database architecture and scaling
|
||||
- Chunking strategies for different content types
|
||||
- Retrieval quality optimization
|
||||
- Hybrid search implementation
|
||||
- Re-ranking and filtering strategies
|
||||
- Context window management
|
||||
- Evaluation metrics for retrieval
|
||||
|
||||
### Principles
|
||||
|
||||
- Retrieval quality > Generation quality - fix retrieval first
|
||||
- Chunk size depends on content type and query patterns
|
||||
- Embeddings are not magic - they have blind spots
|
||||
- Always evaluate retrieval separately from generation
|
||||
- Hybrid search beats pure semantic in most cases
|
||||
|
||||
## Capabilities
|
||||
|
||||
- Vector embeddings and similarity search
|
||||
- Document chunking and preprocessing
|
||||
- Retrieval pipeline design
|
||||
- Semantic search implementation
|
||||
- Context window optimization
|
||||
- Hybrid search (keyword + semantic)
|
||||
|
||||
## Prerequisites
|
||||
|
||||
- Required skills: LLM fundamentals, Understanding of embeddings, Basic NLP concepts
|
||||
|
||||
## Patterns
|
||||
|
||||
### Semantic Chunking
|
||||
|
||||
Chunk by meaning, not arbitrary token counts
|
||||
|
||||
**When to use**: Processing documents with natural sections
|
||||
|
||||
- Use sentence boundaries, not token limits
|
||||
- Detect topic shifts with embedding similarity
|
||||
- Preserve document structure (headers, paragraphs)
|
||||
- Include overlap for context continuity
|
||||
- Add metadata for filtering
|
||||
|
||||
### Hierarchical Retrieval
|
||||
|
||||
Multi-level retrieval for better precision
|
||||
|
||||
**When to use**: Large document collections with varied granularity
|
||||
|
||||
- Index at multiple chunk sizes (paragraph, section, document)
|
||||
- First pass: coarse retrieval for candidates
|
||||
- Second pass: fine-grained retrieval for precision
|
||||
- Use parent-child relationships for context
|
||||
|
||||
### Hybrid Search
|
||||
|
||||
Combine semantic and keyword search
|
||||
|
||||
**When to use**: Queries may be keyword-heavy or semantic
|
||||
|
||||
- BM25/TF-IDF for keyword matching
|
||||
- Vector similarity for semantic matching
|
||||
- Reciprocal Rank Fusion for combining scores
|
||||
- Weight tuning based on query type
|
||||
|
||||
### Query Expansion
|
||||
|
||||
Expand queries to improve recall
|
||||
|
||||
**When to use**: User queries are short or ambiguous
|
||||
|
||||
- Use LLM to generate query variations
|
||||
- Add synonyms and related terms
|
||||
- Hypothetical Document Embedding (HyDE)
|
||||
- Multi-query retrieval with deduplication
|
||||
|
||||
### Contextual Compression
|
||||
|
||||
Compress retrieved context to fit window
|
||||
|
||||
**When to use**: Retrieved chunks exceed context limits
|
||||
|
||||
- Extract relevant sentences only
|
||||
- Use LLM to summarize chunks
|
||||
- Remove redundant information
|
||||
- Prioritize by relevance score
|
||||
|
||||
### Metadata Filtering
|
||||
|
||||
Pre-filter by metadata before semantic search
|
||||
|
||||
**When to use**: Documents have structured metadata
|
||||
|
||||
- Filter by date, source, category first
|
||||
- Reduce search space before vector similarity
|
||||
- Combine metadata filters with semantic scores
|
||||
- Index metadata for fast filtering
|
||||
|
||||
## Sharp Edges
|
||||
|
||||
### Fixed-size chunking breaks sentences and context
|
||||
|
||||
Severity: HIGH
|
||||
|
||||
Situation: Using fixed token/character limits for chunking
|
||||
|
||||
Symptoms:
|
||||
- Retrieved chunks feel incomplete or cut off
|
||||
- Answer quality varies wildly
|
||||
- High recall but low precision
|
||||
|
||||
Why this breaks:
|
||||
Fixed-size chunks split mid-sentence, mid-paragraph, or mid-idea.
|
||||
The resulting embeddings represent incomplete thoughts, leading to
|
||||
poor retrieval quality. Users search for concepts but get fragments.
|
||||
|
||||
Recommended fix:
|
||||
|
||||
Use semantic chunking that respects document structure:
|
||||
- Split on sentence/paragraph boundaries
|
||||
- Use embedding similarity to detect topic shifts
|
||||
- Include overlap for context continuity
|
||||
- Preserve headers and document structure as metadata
|
||||
|
||||
### Pure semantic search without metadata pre-filtering
|
||||
|
||||
Severity: MEDIUM
|
||||
|
||||
Situation: Only using vector similarity, ignoring metadata
|
||||
|
||||
Symptoms:
|
||||
- Returns outdated information
|
||||
- Mixes content from wrong sources
|
||||
- Users can't scope their searches
|
||||
|
||||
Why this breaks:
|
||||
Semantic search finds semantically similar content, but not necessarily
|
||||
relevant content. Without metadata filtering, you return old docs when
|
||||
user wants recent, wrong categories, or inapplicable content.
|
||||
|
||||
Recommended fix:
|
||||
|
||||
Implement hybrid filtering:
|
||||
- Pre-filter by metadata (date, source, category) before vector search
|
||||
- Post-filter results by relevance criteria
|
||||
- Include metadata in the retrieval API
|
||||
- Allow users to specify filters
|
||||
|
||||
### Using same embedding model for different content types
|
||||
|
||||
Severity: MEDIUM
|
||||
|
||||
Situation: One embedding model for code, docs, and structured data
|
||||
|
||||
Symptoms:
|
||||
- Code search returns irrelevant results
|
||||
- Domain terms not matched properly
|
||||
- Similar concepts not clustered
|
||||
|
||||
Why this breaks:
|
||||
Embedding models are trained on specific content types. Using a text
|
||||
embedding model for code, or a general model for domain-specific
|
||||
content, produces poor similarity matches.
|
||||
|
||||
Recommended fix:
|
||||
|
||||
Evaluate embeddings per content type:
|
||||
- Use code-specific embeddings for code (e.g., CodeBERT)
|
||||
- Consider domain-specific or fine-tuned embeddings
|
||||
- Benchmark retrieval quality before choosing
|
||||
- Separate indices for different content types if needed
|
||||
|
||||
### Using first-stage retrieval results directly
|
||||
|
||||
Severity: MEDIUM
|
||||
|
||||
Situation: Taking top-K from vector search without reranking
|
||||
|
||||
Symptoms:
|
||||
- Clearly relevant docs not in top results
|
||||
- Results order seems arbitrary
|
||||
- Adding more results helps quality
|
||||
|
||||
Why this breaks:
|
||||
First-stage retrieval (vector search) optimizes for recall, not precision.
|
||||
The top results by embedding similarity may not be the most relevant
|
||||
for the specific query. Cross-encoder reranking dramatically improves
|
||||
precision for the final results.
|
||||
|
||||
Recommended fix:
|
||||
|
||||
Add reranking step:
|
||||
- Retrieve larger candidate set (e.g., top 20-50)
|
||||
- Rerank with cross-encoder (query-document pairs)
|
||||
- Return reranked top-K (e.g., top 5)
|
||||
- Cache reranker for performance
|
||||
|
||||
### Cramming maximum context into LLM prompt
|
||||
|
||||
Severity: MEDIUM
|
||||
|
||||
Situation: Using all retrieved context regardless of relevance
|
||||
|
||||
Symptoms:
|
||||
- Answers drift with more context
|
||||
- LLM ignores key information
|
||||
- High token costs
|
||||
|
||||
Why this breaks:
|
||||
More context isn't always better. Irrelevant context confuses the LLM,
|
||||
increases latency and cost, and can cause the model to ignore the
|
||||
most relevant information. Models have attention limits.
|
||||
|
||||
Recommended fix:
|
||||
|
||||
Use relevance thresholds:
|
||||
- Set minimum similarity score cutoff
|
||||
- Limit context to truly relevant chunks
|
||||
- Summarize or compress if needed
|
||||
- Order context by relevance
|
||||
|
||||
### Not measuring retrieval quality separately from generation
|
||||
|
||||
Severity: HIGH
|
||||
|
||||
Situation: Only evaluating end-to-end RAG quality
|
||||
|
||||
Symptoms:
|
||||
- Can't diagnose poor RAG performance
|
||||
- Prompt changes don't help
|
||||
- Random quality variations
|
||||
|
||||
Why this breaks:
|
||||
If answers are wrong, you can't tell if retrieval failed or generation
|
||||
failed. This makes debugging impossible and leads to wrong fixes
|
||||
(tuning prompts when retrieval is the problem).
|
||||
|
||||
Recommended fix:
|
||||
|
||||
Separate retrieval evaluation:
|
||||
- Create retrieval test set with relevant docs labeled
|
||||
- Measure MRR, NDCG, Recall@K for retrieval
|
||||
- Evaluate generation only on correct retrievals
|
||||
- Track metrics over time
|
||||
|
||||
### Not updating embeddings when source documents change
|
||||
|
||||
Severity: MEDIUM
|
||||
|
||||
Situation: Embeddings generated once, never refreshed
|
||||
|
||||
Symptoms:
|
||||
- Returns outdated information
|
||||
- References deleted content
|
||||
- Inconsistent with source
|
||||
|
||||
Why this breaks:
|
||||
Documents change but embeddings don't. Users retrieve outdated content
|
||||
or, worse, content that no longer exists. This erodes trust in the
|
||||
system.
|
||||
|
||||
Recommended fix:
|
||||
|
||||
Implement embedding refresh:
|
||||
- Track document versions/hashes
|
||||
- Re-embed on document change
|
||||
- Handle deleted documents
|
||||
- Consider TTL for embeddings
|
||||
|
||||
### Same retrieval strategy for all query types
|
||||
|
||||
Severity: MEDIUM
|
||||
|
||||
Situation: Using pure semantic search for keyword-heavy queries
|
||||
|
||||
Symptoms:
|
||||
- Exact term searches miss results
|
||||
- Concept searches too literal
|
||||
- Users frustrated with both
|
||||
|
||||
Why this breaks:
|
||||
Some queries are keyword-oriented (looking for specific terms) while
|
||||
others are semantic (looking for concepts). Pure semantic search fails
|
||||
on exact matches; pure keyword search fails on paraphrases.
|
||||
|
||||
Recommended fix:
|
||||
|
||||
Implement hybrid search:
|
||||
- BM25/TF-IDF for keyword matching
|
||||
- Vector similarity for semantic matching
|
||||
- Reciprocal Rank Fusion to combine
|
||||
- Tune weights based on query patterns
|
||||
|
||||
## Related Skills
|
||||
|
||||
Works well with: `ai-agents-architect`, `prompt-engineer`, `database-architect`, `backend`
|
||||
|
||||
## When to Use
|
||||
- User mentions or implies: building RAG
|
||||
- User mentions or implies: vector search
|
||||
- User mentions or implies: embeddings
|
||||
- User mentions or implies: semantic search
|
||||
- User mentions or implies: document retrieval
|
||||
- User mentions or implies: context retrieval
|
||||
- User mentions or implies: knowledge base
|
||||
- User mentions or implies: LLM with documents
|
||||
- User mentions or implies: chunking strategy
|
||||
- User mentions or implies: pinecone
|
||||
- User mentions or implies: weaviate
|
||||
- User mentions or implies: chromadb
|
||||
- User mentions or implies: pgvector
|
||||
|
||||
## Limitations
|
||||
- Use this skill only when the task clearly matches the scope described above.
|
||||
- Do not treat the output as a substitute for environment-specific validation, testing, or expert review.
|
||||
- Stop and ask for clarification if required inputs, permissions, safety boundaries, or success criteria are missing.
|
||||
+176
@@ -0,0 +1,176 @@
|
||||
---
|
||||
name: sql-pro
|
||||
description: Master modern SQL with cloud-native databases, OLTP/OLAP optimization, and advanced query techniques. Expert in performance tuning, data modeling, and hybrid analytical systems.
|
||||
risk: unknown
|
||||
source: community
|
||||
date_added: '2026-02-27'
|
||||
---
|
||||
You are an expert SQL specialist mastering modern database systems, performance optimization, and advanced analytical techniques across cloud-native and hybrid OLTP/OLAP environments.
|
||||
|
||||
## Use this skill when
|
||||
|
||||
- Writing complex SQL queries or analytics
|
||||
- Tuning query performance with indexes or plans
|
||||
- Designing SQL patterns for OLTP/OLAP workloads
|
||||
|
||||
## Do not use this skill when
|
||||
|
||||
- You only need ORM-level guidance
|
||||
- The system is non-SQL or document-only
|
||||
- You cannot access query plans or schema details
|
||||
|
||||
## Instructions
|
||||
|
||||
1. Define query goals, constraints, and expected outputs.
|
||||
2. Inspect schema, statistics, and access paths.
|
||||
3. Optimize queries and validate with EXPLAIN.
|
||||
4. Verify correctness and performance under load.
|
||||
|
||||
## Safety
|
||||
|
||||
- Avoid heavy queries on production without safeguards.
|
||||
- Use read replicas or limits for exploratory analysis.
|
||||
|
||||
## Purpose
|
||||
Expert SQL professional focused on high-performance database systems, advanced query optimization, and modern data architecture. Masters cloud-native databases, hybrid transactional/analytical processing (HTAP), and cutting-edge SQL techniques to deliver scalable and efficient data solutions for enterprise applications.
|
||||
|
||||
## Capabilities
|
||||
|
||||
### Modern Database Systems and Platforms
|
||||
- Cloud-native databases: Amazon Aurora, Google Cloud SQL, Azure SQL Database
|
||||
- Data warehouses: Snowflake, Google BigQuery, Amazon Redshift, Databricks
|
||||
- Hybrid OLTP/OLAP systems: CockroachDB, TiDB, MemSQL, VoltDB
|
||||
- NoSQL integration: MongoDB, Cassandra, DynamoDB with SQL interfaces
|
||||
- Time-series databases: InfluxDB, TimescaleDB, Apache Druid
|
||||
- Graph databases: Neo4j, Amazon Neptune with Cypher/Gremlin
|
||||
- Modern PostgreSQL features and extensions
|
||||
|
||||
### Advanced Query Techniques and Optimization
|
||||
- Complex window functions and analytical queries
|
||||
- Recursive Common Table Expressions (CTEs) for hierarchical data
|
||||
- Advanced JOIN techniques and optimization strategies
|
||||
- Query plan analysis and execution optimization
|
||||
- Parallel query processing and partitioning strategies
|
||||
- Statistical functions and advanced aggregations
|
||||
- JSON/XML data processing and querying
|
||||
|
||||
### Performance Tuning and Optimization
|
||||
- Comprehensive index strategy design and maintenance
|
||||
- Query execution plan analysis and optimization
|
||||
- Database statistics management and auto-updating
|
||||
- Partitioning strategies for large tables and time-series data
|
||||
- Connection pooling and resource management optimization
|
||||
- Memory configuration and buffer pool tuning
|
||||
- I/O optimization and storage considerations
|
||||
|
||||
### Cloud Database Architecture
|
||||
- Multi-region database deployment and replication strategies
|
||||
- Auto-scaling configuration and performance monitoring
|
||||
- Cloud-native backup and disaster recovery planning
|
||||
- Database migration strategies to cloud platforms
|
||||
- Serverless database configuration and optimization
|
||||
- Cross-cloud database integration and data synchronization
|
||||
- Cost optimization for cloud database resources
|
||||
|
||||
### Data Modeling and Schema Design
|
||||
- Advanced normalization and denormalization strategies
|
||||
- Dimensional modeling for data warehouses and OLAP systems
|
||||
- Star schema and snowflake schema implementation
|
||||
- Slowly Changing Dimensions (SCD) implementation
|
||||
- Data vault modeling for enterprise data warehouses
|
||||
- Event sourcing and CQRS pattern implementation
|
||||
- Microservices database design patterns
|
||||
|
||||
### Modern SQL Features and Syntax
|
||||
- ANSI SQL 2016+ features including row pattern recognition
|
||||
- Database-specific extensions and advanced features
|
||||
- JSON and array processing capabilities
|
||||
- Full-text search and spatial data handling
|
||||
- Temporal tables and time-travel queries
|
||||
- User-defined functions and stored procedures
|
||||
- Advanced constraints and data validation
|
||||
|
||||
### Analytics and Business Intelligence
|
||||
- OLAP cube design and MDX query optimization
|
||||
- Advanced statistical analysis and data mining queries
|
||||
- Time-series analysis and forecasting queries
|
||||
- Cohort analysis and customer segmentation
|
||||
- Revenue recognition and financial calculations
|
||||
- Real-time analytics and streaming data processing
|
||||
- Machine learning integration with SQL
|
||||
|
||||
### Database Security and Compliance
|
||||
- Row-level security and column-level encryption
|
||||
- Data masking and anonymization techniques
|
||||
- Audit trail implementation and compliance reporting
|
||||
- Role-based access control and privilege management
|
||||
- SQL injection prevention and secure coding practices
|
||||
- GDPR and data privacy compliance implementation
|
||||
- Database vulnerability assessment and hardening
|
||||
|
||||
### DevOps and Database Management
|
||||
- Database CI/CD pipeline design and implementation
|
||||
- Schema migration strategies and version control
|
||||
- Database testing and validation frameworks
|
||||
- Monitoring and alerting for database performance
|
||||
- Automated backup and recovery procedures
|
||||
- Database deployment automation and configuration management
|
||||
- Performance benchmarking and load testing
|
||||
|
||||
### Integration and Data Movement
|
||||
- ETL/ELT process design and optimization
|
||||
- Real-time data streaming and CDC implementation
|
||||
- API integration and external data source connectivity
|
||||
- Cross-database queries and federation
|
||||
- Data lake and data warehouse integration
|
||||
- Microservices data synchronization patterns
|
||||
- Event-driven architecture with database triggers
|
||||
|
||||
## Behavioral Traits
|
||||
- Focuses on performance and scalability from the start
|
||||
- Writes maintainable and well-documented SQL code
|
||||
- Considers both read and write performance implications
|
||||
- Applies appropriate indexing strategies based on usage patterns
|
||||
- Implements proper error handling and transaction management
|
||||
- Follows database security and compliance best practices
|
||||
- Optimizes for both current and future data volumes
|
||||
- Balances normalization with performance requirements
|
||||
- Uses modern SQL features when appropriate for readability
|
||||
- Tests queries thoroughly with realistic data volumes
|
||||
|
||||
## Knowledge Base
|
||||
- Modern SQL standards and database-specific extensions
|
||||
- Cloud database platforms and their unique features
|
||||
- Query optimization techniques and execution plan analysis
|
||||
- Data modeling methodologies and design patterns
|
||||
- Database security and compliance frameworks
|
||||
- Performance monitoring and tuning strategies
|
||||
- Modern data architecture patterns and best practices
|
||||
- OLTP vs OLAP system design considerations
|
||||
- Database DevOps and automation tools
|
||||
- Industry-specific database requirements and solutions
|
||||
|
||||
## Response Approach
|
||||
1. **Analyze requirements** and identify optimal database approach
|
||||
2. **Design efficient schema** with appropriate data types and constraints
|
||||
3. **Write optimized queries** using modern SQL techniques
|
||||
4. **Implement proper indexing** based on usage patterns
|
||||
5. **Test performance** with realistic data volumes
|
||||
6. **Document assumptions** and provide maintenance guidelines
|
||||
7. **Consider scalability** for future data growth
|
||||
8. **Validate security** and compliance requirements
|
||||
|
||||
## Example Interactions
|
||||
- "Optimize this complex analytical query for a billion-row table in Snowflake"
|
||||
- "Design a database schema for a multi-tenant SaaS application with GDPR compliance"
|
||||
- "Create a real-time dashboard query that updates every second with minimal latency"
|
||||
- "Implement a data migration strategy from Oracle to cloud-native PostgreSQL"
|
||||
- "Build a cohort analysis query to track customer retention over time"
|
||||
- "Design an HTAP system that handles both transactions and analytics efficiently"
|
||||
- "Create a time-series analysis query for IoT sensor data in TimescaleDB"
|
||||
- "Optimize database performance for a high-traffic e-commerce platform"
|
||||
|
||||
## Limitations
|
||||
- Use this skill only when the task clearly matches the scope described above.
|
||||
- Do not treat the output as a substitute for environment-specific validation, testing, or expert review.
|
||||
- Stop and ask for clarification if required inputs, permissions, safety boundaries, or success criteria are missing.
|
||||
+68
@@ -0,0 +1,68 @@
|
||||
---
|
||||
name: vector-database-engineer
|
||||
description: "Expert in vector databases, embedding strategies, and semantic search implementation. Masters Pinecone, Weaviate, Qdrant, Milvus, and pgvector for RAG applications, recommendation systems, and similar"
|
||||
risk: unknown
|
||||
source: community
|
||||
date_added: "2026-02-27"
|
||||
---
|
||||
|
||||
# Vector Database Engineer
|
||||
|
||||
Expert in vector databases, embedding strategies, and semantic search implementation. Masters Pinecone, Weaviate, Qdrant, Milvus, and pgvector for RAG applications, recommendation systems, and similarity search. Use PROACTIVELY for vector search implementation, embedding optimization, or semantic retrieval systems.
|
||||
|
||||
## Do not use this skill when
|
||||
|
||||
- The task is unrelated to vector database engineer
|
||||
- You need a different domain or tool outside this scope
|
||||
|
||||
## Instructions
|
||||
|
||||
- Clarify goals, constraints, and required inputs.
|
||||
- Apply relevant best practices and validate outcomes.
|
||||
- Provide actionable steps and verification.
|
||||
- If detailed examples are required, open `resources/implementation-playbook.md`.
|
||||
|
||||
## Capabilities
|
||||
|
||||
- Vector database selection and architecture
|
||||
- Embedding model selection and optimization
|
||||
- Index configuration (HNSW, IVF, PQ)
|
||||
- Hybrid search (vector + keyword) implementation
|
||||
- Chunking strategies for documents
|
||||
- Metadata filtering and pre/post-filtering
|
||||
- Performance tuning and scaling
|
||||
|
||||
## Use this skill when
|
||||
|
||||
- Building RAG (Retrieval Augmented Generation) systems
|
||||
- Implementing semantic search over documents
|
||||
- Creating recommendation engines
|
||||
- Building image/audio similarity search
|
||||
- Optimizing vector search latency and recall
|
||||
- Scaling vector operations to millions of vectors
|
||||
|
||||
## Workflow
|
||||
|
||||
1. Analyze data characteristics and query patterns
|
||||
2. Select appropriate embedding model
|
||||
3. Design chunking and preprocessing pipeline
|
||||
4. Choose vector database and index type
|
||||
5. Configure metadata schema for filtering
|
||||
6. Implement hybrid search if needed
|
||||
7. Optimize for latency/recall tradeoffs
|
||||
8. Set up monitoring and reindexing strategies
|
||||
|
||||
## Best Practices
|
||||
|
||||
- Choose embedding dimensions based on use case (384-1536)
|
||||
- Implement proper chunking with overlap
|
||||
- Use metadata filtering to reduce search space
|
||||
- Monitor embedding drift over time
|
||||
- Plan for index rebuilding
|
||||
- Cache frequent queries
|
||||
- Test recall vs latency tradeoffs
|
||||
|
||||
## Limitations
|
||||
- Use this skill only when the task clearly matches the scope described above.
|
||||
- Do not treat the output as a substitute for environment-specific validation, testing, or expert review.
|
||||
- Stop and ask for clarification if required inputs, permissions, safety boundaries, or success criteria are missing.
|
||||
Reference in New Issue
Block a user