Metrics
Detailed guide to understanding and using metrics for monitoring your DanubeData resources.
Overview
Metrics provide quantitative measurements of resource performance and health over time.
VPS Metrics
CPU Metrics
CPU Usage %
- Current CPU utilization
- Range: 0-100%
- Alert: > 80% sustained
CPU Load Average
- System load over 1, 5, 15 minutes
- Varies by CPU count
- Alert: > CPU count
CPU Steal
- CPU time stolen by hypervisor (should be near 0)
- Alert: > 5%
Memory Metrics
Memory Usage %
- RAM utilization
- Range: 0-100%
- Alert: > 85%
Memory Available
- Free memory plus cache/buffers
- Alert: < 10% of total
Swap Usage
- Swap space used
- Alert: > 0 (indicates memory pressure)
Disk Metrics
Disk I/O
- Read/write MB/s
- IOPS (operations per second)
Disk Usage %
- Storage consumption
- Alert: > 85%
Network Metrics
Network In/Out
- Bandwidth usage (MB/s)
- Track against allocation
Network Packets
- Packets per second
- Useful for diagnosing issues
Database Metrics
Performance Metrics
Query Time
- Average query execution time
- Alert: > 100ms average
Slow Queries
- Queries exceeding threshold
- Alert: > 10/minute
Throughput
- Queries per second
- Monitor for capacity planning
Connection Metrics
Active Connections
- Current client connections
- Alert: > 80% of max_connections
Connection Rate
- New connections per second
- Alert: Sudden spikes
Cache Metrics
Buffer Cache Hit Rate
- Percentage of queries served from cache
- Target: > 99%
- Alert: < 95%
Cache Size
- Memory used for query cache
- Monitor for sizing
Replication Metrics
Replication Lag
- Delay between primary and replica
- Alert: > 5 seconds
Replication Status
- Connected/Disconnected status
- Alert: Disconnected
Database load (Performance tab, PostgreSQL)
Managed PostgreSQL instances also have a Performance tab. It answers a different question than the metrics above: not how busy is the server, but what are sessions waiting on, and who is causing it.
Database load
- Average active sessions (AAS) per bucket, sampled every 30 seconds while the cluster runs
- Any range from five minutes to 15 days: quick ranges, or Custom for relative presets (minutes, hours, days, weeks), a free duration, or an absolute start and end in your local time
- Stacked bars or lines; one-minute buckets by default (coarser on long ranges); up to eight series, the rest fold into Other
- Auto-refresh every minute by default, or 30 seconds, 5 or 15 minutes, or off; remembered per browser
Sliced by
- Waits (CPU, disk reads, WAL writes, locks, client waits…), users, applications, databases, hosts, session types, or SQL (the statements themselves)
- Hosts you own are named for you: your VPS instances, your serverless containers and managed apps, other workloads in your namespace, and the connection pooler
Top SQL
- The statements that carried the most load over the range: text, load in AAS, share of all load, and the wait events behind it
- Sampled every 5 seconds on the primary by the query agent SQL Studio uses; fills in a minute or two after the agent starts for your project
- The statement text is kept as first seen, up to 1 KiB
Capacity
- The instance's vCPU count, drawn as a dashed line when load is near it
- Sustained load above that line means sessions are queueing for a CPU: look at the heaviest wait or user, or scale up
Top tables and lock analysis
- The top values of the selected slice over the whole range, with their share of the load
- Lock analysis lists sessions blocked on a lock right now and the session holding it, with both queries. It runs live through the query agent SQL Studio uses; the first start for a project takes about a minute
Selecting a node on a cluster with replicas shows that node alone; the SQL slice and Top SQL always describe the primary. An idle database shows a zeroed chart, not an error.
Cache Metrics
Applies to Redis, Valkey, and Dragonfly instances.
Memory Metrics
Memory Usage
- Current RAM consumption
- Alert: > 90% of allocated
Memory Fragmentation
- Ratio of RSS to used memory
- Alert: > 1.5 (consider restart)
Evicted Keys
- Keys removed due to memory pressure
- Alert: > 100/second
Performance Metrics
Hit Rate
- Cache hit percentage
- Target: > 90%
- Alert: < 80%
Operations/Sec
- Commands processed per second
- Monitor for capacity
Latency
- Average command execution time
- Alert: > 10ms
Connection Metrics
Connected Clients
- Active connections
- Alert: > 80% of max
Blocked Clients
- Clients waiting on blocking operations
- Alert: Sustained blocked clients
Metric History
Metrics history is available for up to 7 days per instance. The chart resolution adjusts to the range you pick — finer detail for recent windows, coarser for the full 7 days. Object storage buckets keep longer per-bucket trends of their own; see the Object Storage docs.
API Access
Cache and database metrics are also available over the public API, so you can pull them into your own dashboards or scripts. See the API Reference for authentication and available endpoints.
Best Practices
- Regular Monitoring: Check metrics daily
- Set Baselines: Know normal values
- Correlate Metrics: Look at multiple metrics together
- Trend Analysis: Watch for gradual changes
- Alert Configuration: Set meaningful thresholds