ai

LLM Quantization GGUF กับ Container

LLM Quantization GGUF กับ Container

LLM Quantization GGUF

LLM Quantization GGUF กับ Container

LLM Quantization ลดขนาด Model FP32 FP16 เป็น INT8 INT4 ลดขนาด 2-4 เท่า ใช้ RAM น้อยลง Inference เร็วขึ้น รันบน Consumer Hardware

เนื้อหาเกี่ยวข้อง — machine learning algorithm คือ

GGUF GPT-Generated Unified Format llama.cpp Model Weights Tokenizer Metadata ไฟล์เดียว Quantization Q2_K ถึง Q8_0 รันบน CPU ไม่ต้อง GPU

เนื้อหาเกี่ยวข้อง — ทำความเข้าใจ Vector Database Pinecone Message Queue Design —

Quant LevelBitsSize (7B)RAMQuality
Q8_08-bit7.2 GB9.7 GBดีมาก ใกล้ FP16
Q6_K6-bit5.5 GB8.0 GBดีมาก
Q5_K_M5-bit4.8 GB7.3 GBดี (แนะนำ)
Q4_K_M4-bit4.1 GB6.6 GBดี
Q3_K_M3-bit3.3 GB5.8 GBพอใช้
Q2_K2-bit2.7 GB5.2 GBลดลงมาก

Docker Container

LLM Quantization GGUF กับ Container
# Dockerfile — llama.cpp Server Container
# FROM ubuntu:22.04 AS builder
# RUN apt-get update && apt-get install -y build-essential git cmake
# RUN git clone https://github.com/ggerganov/llama.cpp /llama.cpp
# WORKDIR /llama.cpp
# RUN cmake -B build -DLLAMA_SERVER=ON && cmake --build build -j$(nproc)
#
# FROM ubuntu:22.04
# COPY --from=builder /llama.cpp/build/bin/llama-server /usr/local/bin/
# RUN mkdir /models
# EXPOSE 8080
# ENTRYPOINT ["llama-server"]
# CMD ["-m", "/models/model.gguf", "--host", "0.0.0.0", "--port", "8080", "-c", "4096"]

# docker-compose.yml
# version: "3.8"
# services:
#   llm-server:
#     build: .
#     ports:
#       - "8080:8080"
#     volumes:
#       - ./models:/models
#     environment:
#       - MODEL_PATH=/models/llama-2-7b-Q5_K_M.gguf
#       - CONTEXT_SIZE=4096
#       - N_GPU_LAYERS=35
#     deploy:
#       resources:
#         limits:
#           memory: 8G
#           cpus: "4"
#         reservations:
#           memory: 6G

# Kubernetes Deployment
# apiVersion: apps/v1
# kind: Deployment
# metadata:
#   name: llm-server
# spec:
#   replicas: 3
#   selector:
#     matchLabels:
#       app: llm-server
#   template:
#     spec:
#       containers:
#       - name: llm
#         image: llama-cpp-server:latest
#         args:
#         - "-m"
#         - "/models/llama-2-7b-Q5_K_M.gguf"
#         - "--host"
#         - "0.0.0.0"
#         - "--port"
#         - "8080"
#         - "-c"
#         - "4096"
#         ports:
#         - containerPort: 8080
#         resources:
#           requests:
#             memory: 6Gi
#             cpu: 2
#           limits:
#             memory: 8Gi
#             cpu: 4
#         volumeMounts:
#         - name: models
#           mountPath: /models
#         readinessProbe:
#           httpGet:
#             path: /health
#             port: 8080
#           initialDelaySeconds: 30
#       volumes:
#       - name: models
#         persistentVolumeClaim:
#           claimName: model-pvc

container_configs = {
    "Small (7B Q4)": {"ram": "6-8 GB", "cpu": "2-4 cores", "replicas": "3-5"},
    "Medium (13B Q4)": {"ram": "10-12 GB", "cpu": "4-8 cores", "replicas": "2-3"},
    "Large (70B Q4)": {"ram": "40-48 GB", "cpu": "8-16 cores", "replicas": "1-2"},
    "GPU (7B FP16)": {"ram": "16 GB VRAM", "cpu": "4 cores", "replicas": "3-5"},
}

print("Container Configurations:")
for config, specs in container_configs.items():
    print(f"\n  [{config}]")
    for key, value in specs.items():
        print(f"    {key}: {value}")

Production Architecture

# production.py — Production LLM Architecture
architecture = {
    "Load Balancer": "Nginx/Traefik กระจาย Requests ไป LLM Pods",
    "LLM Pods": "llama.cpp Server 3-5 Replicas + HPA",
    "Model Storage": "PVC (NFS/EBS) แชร์ GGUF Files ข้าม Pods",
    "Cache": "Redis Cache ผลลัพธ์ที่เคยถามแล้ว (Semantic Cache)",
    "Queue": "Redis/Kafka Queue Requests ป้องกัน Overload",
    "Monitoring": "Prometheus + Grafana: Tokens/s, Latency, Memory",
    "Rate Limiting": "จำกัด Requests/min ต่อ User",
}

print("Production LLM Architecture:")
for component, desc in architecture.items():
    print(f"  [{component}]")
    print(f"    {desc}")

# Popular GGUF Models
models = {
    "Llama-3-8B": {"size_q5": "5.5 GB", "context": "8K", "license": "Meta"},
    "Mistral-7B": {"size_q5": "4.8 GB", "context": "32K", "license": "Apache 2.0"},
    "Phi-3-mini": {"size_q5": "2.4 GB", "context": "128K", "license": "MIT"},
    "Gemma-2-9B": {"size_q5": "6.1 GB", "context": "8K", "license": "Google"},
    "Qwen2-7B": {"size_q5": "4.8 GB", "context": "128K", "license": "Apache 2.0"},
    "CodeLlama-7B": {"size_q5": "4.8 GB", "context": "16K", "license": "Meta"},
}

print(f"\n\nPopular GGUF Models (Q5_K_M):")
for model, info in models.items():
    print(f"  {model}: {info['size_q5']} | Context: {info['context']} | License: {info['license']}")

Best Practices

  • Q5_K_M: แนะนำสำหรับใช้ทั่วไป สมดุล Quality กับ Size
  • PVC: ใช้ PVC เก็บ Model Files แชร์ข้าม Pods ไม่ต้อง Download ซ้ำ
  • Semantic Cache: Cache คำถามที่คล้ายกัน ลด Inference 80%+
  • HPA: Scale ตาม CPU/Memory ไม่ใช่ Request Count
  • Health Check: ตรวจ /health ทุก 30 วินาที Restart ถ้า OOM
  • Context Length: จำกัด Context Length ลด RAM ใช้

LLM Quantization คืออะไร

ลดขนาด LLM แปลง FP32 FP16 เป็น INT8 INT4 ลดขนาด 2-4 เท่า RAM น้อยลง Inference เร็วขึ้น รันบน Consumer Hardware Llama 70B 140GB เหลือ 40GB

แนะนำเพิ่มเติม — สัญญาณเทรดรายวัน XM Signal

เนื้อหาเกี่ยวข้อง — อ่านต่อ: DNS over TLS Machine Learning Pipeline

XM Legend · เทรดเดอร์ & ผู้สอน Forex 13 ปี

ผู้ก่อตั้ง SiamCafe ตั้งแต่ปี 1997 · เทรดเดอร์สาย Forex มากกว่า 13 ปี ได้รับการยกย่องเป็น XM Legend · แบ่งปันความรู้ Forex, ไอที, AI และการเทรด จากประสบการณ์จริงในตลาดจริง