ESC

开始输入,可搜索发票、服务、域名、工单,以及 更多...

搜索... Ctrl+K
Linux 服务器

Linux 服务器监控教程:查看 CPU、内存、磁盘 I/O 与带宽占用(top、iostat、iftop、vnstat)

7 个步骤 10 分钟阅读 3 次阅读 0
本文目录

网站变慢、SSH 卡顿、流量突然暴涨时,第一步是弄清楚瓶颈在哪里。本文整理一套实用的 Linux 服务器监控方法:用 top/htop 查看 CPU 与进程,free 查看内存,iostat、iotop 定位磁盘 I/O,iftop、nload、vnstat 查看带宽占用,ss 找出占用连接的 IP,最后写一个 cron 脚本实现简单告警。适用于 Debian/Ubuntu 与 Rocky Linux/AlmaLinux 的 VPS 和独立服务器。

步骤 1:安装常用监控工具

top、free、ss 系统自带,其余工具需要安装;Rocky/AlmaLinux 需先启用 EPEL。后面命令中的网卡名以 eth0 为例,请先用 ip -br addr 确认实际名称。

# Debian / Ubuntu
apt install -y htop sysstat iotop iftop nload vnstat nethogs
# Rocky Linux / AlmaLinux (most tools are in EPEL)
dnf install -y epel-release
dnf install -y htop sysstat iotop iftop nload vnstat nethogs

ip -br addr          # find your network interface name (eth0, ens3, ...)

步骤 2:查看 CPU 占用与负载(top/htop)

load average 的三个数字分别是 1、5、15 分钟平均负载,长期高于 CPU 核数说明处理不过来。在 top 中重点看三项:us 高是程序本身耗 CPU;wa 高说明在等磁盘,应转去看 I/O;VPS 上 st(steal)持续较高表示宿主资源争用,可提交工单让技术支持协助检查。

nproc                # number of CPU cores
uptime               # load average for 1, 5 and 15 minutes
top                  # 1 = per-core view, P = sort by CPU, M = sort by memory, q = quit
htop                 # F6 to choose the sort column, F9 to send a signal
ps aux --sort=-%cpu | head -n 10

步骤 3:查看内存使用与 OOM

判断内存是否紧张请看 free -h 的 available 列,而不是 free 列——Linux 会把空闲内存用作缓存,free 小是正常的。如果 vmstat 中 si/so 持续不为 0,说明正在频繁使用 Swap;进程莫名消失时,用 dmesg 检查是否被 OOM Killer 结束。内存不足时可先添加 Swap(参见“Swap”教程)或升级配置。

free -h
vmstat 1 5                                   # si/so > 0 continuously = swapping
ps aux --sort=-%mem | head -n 10
dmesg -T | grep -iE "out of memory|killed process"
journalctl -k --since "1 day ago" | grep -i oom

步骤 4:排查磁盘 I/O 与磁盘空间

iostat 中 %util 接近 100%、await 明显升高,说明磁盘已经忙不过来;再用 iotop 找出读写最多的进程,常见的有数据库、日志写入、备份任务。sysstat 开启后会每 10 分钟记录一次,可用 sar 回看过去的数据。空间不足的处理见“磁盘空间不足”教程。

iostat -xz 1 5        # watch %util, r_await / w_await, rkB/s and wkB/s
iotop -oPa            # only processes doing I/O, accumulated totals
df -h                 # space usage
df -i                 # inode usage

# Keep history with sysstat, then read it with sar
systemctl enable --now sysstat
sar -u                # CPU today
sar -d -p             # disks today

步骤 5:查看实时带宽与流量统计

nload 显示网卡实时进出速率;iftop 按远端 IP 列出流量,能快速看出是谁在大量下载;nethogs 则按进程统计。vnstat 在后台按日、按月累计流量,适合核对带宽或流量使用情况。如果出站流量异常且你没有对应业务,需警惕服务器被入侵后对外发包。

ip -s link show eth0          # total RX/TX bytes since boot
nload eth0                    # live in/out graph
iftop -i eth0 -nNP            # live traffic per remote host and port
nethogs eth0                  # live traffic per process

systemctl enable --now vnstat
vnstat -i eth0 -l             # live rate
vnstat -d                     # daily totals
vnstat -m                     # monthly totals

步骤 6:用 ss 找出占用连接与带宽的来源

ss 是 netstat 的替代品,速度更快。下面的命令统计每个远端 IP 的已建立连接数,可快速发现爬虫或 CC 攻击来源,再结合防火墙或 Fail2ban 进行封禁。

ss -s                                              # connection summary
ss -tunp | head -n 30                              # connections with process names
ss -lntup                                          # listening ports

# Top 10 remote IPs by number of established TCP connections
ss -Htn state established | awk '{sub(/:[0-9]+$/,"",$4); print $4}' \
  | sort | uniq -c | sort -rn | head -n 10

# Connections to your web server only
ss -Htn state established '( sport = :443 )' | wc -l

步骤 7:用 cron 脚本实现简单告警

没有部署 Zabbix、Prometheus 等监控系统时,可用一个小脚本每 5 分钟检查负载、可用内存和磁盘使用率,超过阈值就写日志并推送到 Webhook(Telegram 机器人、Slack、企业微信、钉钉等均可,按各平台要求调整 JSON 格式)。

cat > /usr/local/bin/check-health.sh <<'EOF'
#!/bin/bash
LOAD_MAX=$(nproc)          # alert when 1-min load > number of cores
MEM_MIN_MB=200             # alert when available memory < 200 MB
DISK_MAX=90                # alert when a filesystem is >= 90 % full
WEBHOOK="https://hooks.example.com/your-webhook"
HOST=$(hostname)
MSG=""

LOAD=$(cut -d' ' -f1 /proc/loadavg)
if awk -v l="$LOAD" -v m="$LOAD_MAX" 'BEGIN{exit !(l>m)}'; then
  MSG+="load $LOAD > $LOAD_MAX; "
fi

AVAIL=$(free -m | awk '/^Mem:/{print $7}')
if [ "$AVAIL" -lt "$MEM_MIN_MB" ]; then MSG+="available memory ${AVAIL}MB; "; fi

while read -r USE MNT; do
  if [ "${USE%\%}" -ge "$DISK_MAX" ]; then MSG+="disk $MNT at $USE; "; fi
done < <(df -P -x tmpfs -x devtmpfs | awk 'NR>1{print $5, $6}')

if [ -n "$MSG" ]; then
  echo "$(date '+%F %T') $HOST $MSG" >> /var/log/health-alert.log
  curl -s -m 10 -H 'Content-Type: application/json' \
    -d "{\"text\":\"[$HOST] $MSG\"}" "$WEBHOOK" > /dev/null
fi
EOF
chmod +x /usr/local/bin/check-health.sh
/usr/local/bin/check-health.sh; tail /var/log/health-alert.log

# Run every 5 minutes: crontab -e
*/5 * * * * /usr/local/bin/check-health.sh
告警脚本只是兜底手段。业务量大的服务器建议部署 Netdata、Zabbix 或 Prometheus + Grafana,并从另一台服务器做外部可用性探测,这样整机宕机时也能收到通知。

常见问题

top 里 CPU 不高,但负载很高?

负载包含等待磁盘 I/O 的进程(D 状态)。请查看 wa 值和 iostat,多数是磁盘繁忙或网络存储响应慢导致。

vnstat 显示 “no data available”?

vnstat 刚安装时还没有数据,需要服务运行几分钟后才会出现统计;同时确认 vnstat --iflist 中有你要统计的网卡。

如何找到占用带宽最多的进程?

用 nethogs eth0 按进程查看实时流量,再用 ss -tunp 查看该进程的连接对象,即可判断是正常业务还是异常程序。

如果按以上步骤操作后问题仍未解决,请提交工单联系 IMIDC 7×24 技术支持,并附上服务器 IP、系统版本、执行过的命令和完整报错信息,方便工程师快速定位。

这篇文章有帮助吗?

相关教程