ESC

開始輸入,可搜尋發票、服務、域名、工單,以及 更多...

搜尋... Ctrl+K
Linux 伺服器

Linux 伺服器監控教程:檢視 CPU、記憶體、磁碟 I/O 與頻寬佔用(top、iostat、iftop、vnstat)

7 個步驟 10 分鐘閱讀 15 次閱讀 0
本文目錄

網站變慢、SSH 卡頓、流量突然暴漲時,第一步是弄清楚瓶頸在哪裡。本文整理一套實用的 Linux 伺服器監控方法:用 top/htop 檢視 CPU 與程序,free 檢視記憶體,iostat、iotop 定位磁碟 I/O,iftop、nload、vnstat 檢視頻寬佔用,ss 找出佔用連線的 IP,最後寫一個 cron 指令碼實現簡單告警。適用於 Debian/Ubuntu 與 Rocky Linux/AlmaLinux 的 VPS 和獨立伺服器。

步驟 1:安裝常用監控工具

top、free、ss 系統自帶,其餘工具需要安裝;Rocky/AlmaLinux 需先啟用 EPEL。後面命令中的網絡卡名以 eth0 為例,請先用 ip -br addr 確認實際名稱。

# Debian / Ubuntu
apt install -y htop sysstat iotop iftop nload vnstat nethogs
# Rocky Linux / AlmaLinux (most tools are in EPEL)
dnf install -y epel-release
dnf install -y htop sysstat iotop iftop nload vnstat nethogs

ip -br addr          # find your network interface name (eth0, ens3, ...)

步驟 2:檢視 CPU 佔用與負載(top/htop)

load average 的三個數字分別是 1、5、15 分鐘平均負載,長期高於 CPU 核數說明處理不過來。在 top 中重點看三項:us 高是程式本身耗 CPU;wa 高說明在等磁碟,應轉去看 I/O;VPS 上 st(steal)持續較高表示宿主資源爭用,可提交工單讓技術支援協助檢查。

nproc                # number of CPU cores
uptime               # load average for 1, 5 and 15 minutes
top                  # 1 = per-core view, P = sort by CPU, M = sort by memory, q = quit
htop                 # F6 to choose the sort column, F9 to send a signal
ps aux --sort=-%cpu | head -n 10

步驟 3:檢視記憶體使用與 OOM

判斷記憶體是否緊張請看 free -h 的 available 列,而不是 free 列——Linux 會把空閒記憶體用作快取,free 小是正常的。如果 vmstat 中 si/so 持續不為 0,說明正在頻繁使用 Swap;程序莫名消失時,用 dmesg 檢查是否被 OOM Killer 結束。記憶體不足時可先新增 Swap(參見“Swap”教程)或升級配置。

free -h
vmstat 1 5                                   # si/so > 0 continuously = swapping
ps aux --sort=-%mem | head -n 10
dmesg -T | grep -iE "out of memory|killed process"
journalctl -k --since "1 day ago" | grep -i oom

步驟 4:排查磁碟 I/O 與磁碟空間

iostat 中 %util 接近 100%、await 明顯升高,說明磁碟已經忙不過來;再用 iotop 找出讀寫最多的程序,常見的有資料庫、日誌寫入、備份任務。sysstat 開啟後會每 10 分鐘記錄一次,可用 sar 回看過去的資料。空間不足的處理見“磁碟空間不足”教程。

iostat -xz 1 5        # watch %util, r_await / w_await, rkB/s and wkB/s
iotop -oPa            # only processes doing I/O, accumulated totals
df -h                 # space usage
df -i                 # inode usage

# Keep history with sysstat, then read it with sar
systemctl enable --now sysstat
sar -u                # CPU today
sar -d -p             # disks today

步驟 5:檢視即時頻寬與流量統計

nload 顯示網絡卡即時進出速率;iftop 按遠端 IP 列出流量,能快速看出是誰在大量下載;nethogs 則按程序統計。vnstat 在後臺按日、按月累計流量,適合核對頻寬或流量使用情況。如果出站流量異常且你沒有對應業務,需警惕伺服器被入侵後對外發包。

ip -s link show eth0          # total RX/TX bytes since boot
nload eth0                    # live in/out graph
iftop -i eth0 -nNP            # live traffic per remote host and port
nethogs eth0                  # live traffic per process

systemctl enable --now vnstat
vnstat -i eth0 -l             # live rate
vnstat -d                     # daily totals
vnstat -m                     # monthly totals

步驟 6:用 ss 找出佔用連線與頻寬的來源

ss 是 netstat 的替代品,速度更快。下面的命令統計每個遠端 IP 的已建立連線數,可快速發現爬蟲或 CC 攻擊來源,再結合防火牆或 Fail2ban 進行封禁。

ss -s                                              # connection summary
ss -tunp | head -n 30                              # connections with process names
ss -lntup                                          # listening ports

# Top 10 remote IPs by number of established TCP connections
ss -Htn state established | awk '{sub(/:[0-9]+$/,"",$4); print $4}' \
  | sort | uniq -c | sort -rn | head -n 10

# Connections to your web server only
ss -Htn state established '( sport = :443 )' | wc -l

步驟 7:用 cron 指令碼實現簡單告警

沒有部署 Zabbix、Prometheus 等監控系統時,可用一個小指令碼每 5 分鐘檢查負載、可用記憶體和磁碟使用率,超過閾值就寫日誌並推送到 Webhook(Telegram 機器人、Slack、企業微信、釘釘等均可,按各平台要求調整 JSON 格式)。

cat > /usr/local/bin/check-health.sh <<'EOF'
#!/bin/bash
LOAD_MAX=$(nproc)          # alert when 1-min load > number of cores
MEM_MIN_MB=200             # alert when available memory < 200 MB
DISK_MAX=90                # alert when a filesystem is >= 90 % full
WEBHOOK="https://hooks.example.com/your-webhook"
HOST=$(hostname)
MSG=""

LOAD=$(cut -d' ' -f1 /proc/loadavg)
if awk -v l="$LOAD" -v m="$LOAD_MAX" 'BEGIN{exit !(l>m)}'; then
  MSG+="load $LOAD > $LOAD_MAX; "
fi

AVAIL=$(free -m | awk '/^Mem:/{print $7}')
if [ "$AVAIL" -lt "$MEM_MIN_MB" ]; then MSG+="available memory ${AVAIL}MB; "; fi

while read -r USE MNT; do
  if [ "${USE%\%}" -ge "$DISK_MAX" ]; then MSG+="disk $MNT at $USE; "; fi
done < <(df -P -x tmpfs -x devtmpfs | awk 'NR>1{print $5, $6}')

if [ -n "$MSG" ]; then
  echo "$(date '+%F %T') $HOST $MSG" >> /var/log/health-alert.log
  curl -s -m 10 -H 'Content-Type: application/json' \
    -d "{\"text\":\"[$HOST] $MSG\"}" "$WEBHOOK" > /dev/null
fi
EOF
chmod +x /usr/local/bin/check-health.sh
/usr/local/bin/check-health.sh; tail /var/log/health-alert.log

# Run every 5 minutes: crontab -e
*/5 * * * * /usr/local/bin/check-health.sh
告警指令碼只是兜底手段。業務量大的伺服器建議部署 Netdata、Zabbix 或 Prometheus + Grafana,並從另一台伺服器做外部可用性探測,這樣整機宕機時也能收到通知。

常見問題

top 裡 CPU 不高,但負載很高?

負載包含等待磁碟 I/O 的程序(D 狀態)。請檢視 wa 值和 iostat,多數是磁碟繁忙或網路儲存響應慢導致。

vnstat 顯示 “no data available”?

vnstat 剛安裝時還沒有資料,需要服務執行幾分鐘後才會出現統計;同時確認 vnstat --iflist 中有你要統計的網絡卡。

如何找到佔用頻寬最多的程序?

用 nethogs eth0 按程序檢視即時流量,再用 ss -tunp 檢視該程序的連線物件,即可判斷是正常業務還是異常程式。

如果按以上步驟操作後問題仍未解決,請提交工單聯絡 IMIDC 7×24 技術支援,並附上伺服器 IP、系統版本、執行過的命令和完整報錯資訊,方便工程師快速定位。

這篇文章有幫助嗎?

相關教程