Metrics

Detailed guide to understanding and using metrics for monitoring your DanubeData resources.

Overview

Metrics provide quantitative measurements of resource performance and health over time.

VPS Metrics

CPU Metrics

CPU Usage %

  • Current CPU utilization
  • Range: 0-100%
  • Alert: > 80% sustained

CPU Load Average

  • System load over 1, 5, 15 minutes
  • Varies by CPU count
  • Alert: > CPU count

CPU Steal

  • CPU time stolen by hypervisor (should be near 0)
  • Alert: > 5%

Memory Metrics

Memory Usage %

  • RAM utilization
  • Range: 0-100%
  • Alert: > 85%

Memory Available

  • Free memory plus cache/buffers
  • Alert: < 10% of total

Swap Usage

  • Swap space used
  • Alert: > 0 (indicates memory pressure)

Disk Metrics

Disk I/O

  • Read/write MB/s
  • IOPS (operations per second)

Disk Usage %

  • Storage consumption
  • Alert: > 85%

Network Metrics

Network In/Out

  • Bandwidth usage (MB/s)
  • Track against allocation

Network Packets

  • Packets per second
  • Useful for diagnosing issues

Database Metrics

Performance Metrics

Query Time

  • Average query execution time
  • Alert: > 100ms average

Slow Queries

  • Queries exceeding threshold
  • Alert: > 10/minute

Throughput

  • Queries per second
  • Monitor for capacity planning

Connection Metrics

Active Connections

  • Current client connections
  • Alert: > 80% of max_connections

Connection Rate

  • New connections per second
  • Alert: Sudden spikes

Cache Metrics

Buffer Cache Hit Rate

  • Percentage of queries served from cache
  • Target: > 99%
  • Alert: < 95%

Cache Size

  • Memory used for query cache
  • Monitor for sizing

Replication Metrics

Replication Lag

  • Delay between primary and replica
  • Alert: > 5 seconds

Replication Status

  • Connected/Disconnected status
  • Alert: Disconnected

Database load (Performance tab, PostgreSQL)

Managed PostgreSQL instances also have a Performance tab. It answers a different question than the metrics above: not how busy is the server, but what are sessions waiting on, and who is causing it.

Database load

  • Average active sessions (AAS) per bucket, sampled every 30 seconds while the cluster runs
  • Any range from five minutes to 15 days: quick ranges, or Custom for relative presets (minutes, hours, days, weeks), a free duration, or an absolute start and end in your local time
  • Stacked bars or lines; one-minute buckets by default (coarser on long ranges); up to eight series, the rest fold into Other
  • Auto-refresh every minute by default, or 30 seconds, 5 or 15 minutes, or off; remembered per browser

Sliced by

  • Waits (CPU, disk reads, WAL writes, locks, client waits…), users, applications, databases, hosts, session types, or SQL (the statements themselves)
  • Hosts you own are named for you: your VPS instances, your serverless containers and managed apps, other workloads in your namespace, and the connection pooler

Top SQL

  • The statements that carried the most load over the range: text, load in AAS, share of all load, and the wait events behind it
  • Sampled every 5 seconds on the primary by the query agent SQL Studio uses; fills in a minute or two after the agent starts for your project
  • The statement text is kept as first seen, up to 1 KiB

Capacity

  • The instance's vCPU count, drawn as a dashed line when load is near it
  • Sustained load above that line means sessions are queueing for a CPU: look at the heaviest wait or user, or scale up

Top tables and lock analysis

  • The top values of the selected slice over the whole range, with their share of the load
  • Lock analysis lists sessions blocked on a lock right now and the session holding it, with both queries. It runs live through the query agent SQL Studio uses; the first start for a project takes about a minute

Selecting a node on a cluster with replicas shows that node alone; the SQL slice and Top SQL always describe the primary. An idle database shows a zeroed chart, not an error.

Cache Metrics

Applies to Redis, Valkey, and Dragonfly instances.

Memory Metrics

Memory Usage

  • Current RAM consumption
  • Alert: > 90% of allocated

Memory Fragmentation

  • Ratio of RSS to used memory
  • Alert: > 1.5 (consider restart)

Evicted Keys

  • Keys removed due to memory pressure
  • Alert: > 100/second

Performance Metrics

Hit Rate

  • Cache hit percentage
  • Target: > 90%
  • Alert: < 80%

Operations/Sec

  • Commands processed per second
  • Monitor for capacity

Latency

  • Average command execution time
  • Alert: > 10ms

Connection Metrics

Connected Clients

  • Active connections
  • Alert: > 80% of max

Blocked Clients

  • Clients waiting on blocking operations
  • Alert: Sustained blocked clients

Metric History

Metrics history is available for up to 7 days per instance. The chart resolution adjusts to the range you pick — finer detail for recent windows, coarser for the full 7 days. Object storage buckets keep longer per-bucket trends of their own; see the Object Storage docs.

API Access

Cache and database metrics are also available over the public API, so you can pull them into your own dashboards or scripts. See the API Reference for authentication and available endpoints.

Best Practices

  1. Regular Monitoring: Check metrics daily
  2. Set Baselines: Know normal values
  3. Correlate Metrics: Look at multiple metrics together
  4. Trend Analysis: Watch for gradual changes
  5. Alert Configuration: Set meaningful thresholds