Files
llm-server-setup/README.md
T
2026-09-20 13:43:58 -03:00

3.9 KiB

LLM Server - Deployment Automatizado

Sistema completo de IA agentica con LLMs locales (Phi + DeepSeek 33B) en Ubuntu 22.04.

Requisitos Mínimos

  • CPU: Intel i3-10105 (4 cores)
  • RAM: 16GB DDR4
  • GPU: NVIDIA RTX 3090 (24GB VRAM)
  • SSD: 1TB NVMe (Samsung PM991A o similar)
  • PSU: 850W+
  • SO: Ubuntu 22.04 LTS

Instalación Rápida

# 1. Clonar repositorio
git clone https://github.com/tuuser/llm-server-setup.git
cd llm-server-setup

# 2. Instalar (con sudo)
sudo make install

# 3. Verificar estado
make health-check

# 4. Ver monitoreo
make monitor

Comandos Disponibles

make install              # Instalación completa
make install-quick        # Sin descargar modelos
make update               # Actualizar código
make monitor              # Monitoreo en tiempo real
make health-check         # Verificar estado
make logs                 # Ver logs
make status               # Estado de servicios
make start/stop/restart   # Control de servicios
make backup               # Hacer backup

Estructura de Directorios

llm-server-setup/
├── scripts/             # Scripts de deployment
│ ├── install.sh         # Instalación principal
│ ├── update.sh          # Actualización
│ ├── monitor.sh         # Monitoreo
│ ├── health-check.sh    # Verificación
│ └── backup.sh          # Backup
├── config/              # Configuración
│ └── server.json        # Config del servidor
├── static/              # Web UI
├── workspace/           # Workspace de trabajo
├── logs/                # Archivos de log
├── backups/             # Backups automáticos
├── requirements.txt     # Dependencias Python
├── Makefile             # Automatización
├── .env                 # Variables de entorno
└── README.md            # Este archivo

Servicios

Ollama (LLM Runtime)

  • Puerto: 11434
  • Modelos: Phi (2.7B), DeepSeek Coder (33B)
  • Status: sudo systemctl status ollama

LLM API (FastAPI)

  • Puerto: 8000
  • Health: curl http://localhost:8000/health
  • Status: sudo systemctl status llm-api

Modelos Disponibles

Modelo Tamaño VRAM Velocidad Uso
Phi 2.7B 1.6GB ⚡⚡⚡ Rápido, simple
DeepSeek 33B 33B 18GB ⚡ Potente, razonamiento

Monitoreo

# Monitor en tiempo real
make monitor

# Ver logs
make logs
make logs-api

# Health check
make health-check

API Endpoints

# Health check
curl http://localhost:8000/health

# Listar modelos
curl http://localhost:8000/models

# Generar código
curl -X POST http://localhost:8000/code \
  -H "Content-Type: application/json" \
  -d '{"message":"Hola","model":"phi"}'

Troubleshooting

GPU no se detecta

# Verificar driver
nvidia-smi

# Reinstalar GPU drivers
sudo make install-gpu-only

Servicios no inician

# Ver logs detallados
sudo journalctl -u ollama -f
sudo journalctl -u llm-api -f

# Reiniciar
sudo systemctl restart ollama llm-api

Espacio en disco lleno

# Ver uso
df -h

# Limpiar
make clean

# Hacer backup y restaurar
make backup
rm -rf ~/.ollama/models/*
ollama pull phi:latest

Seguridad

  • SSH habilitado con key-based auth
  • Firewall: Permitir solo puertos necesarios
  • Usuarios: Crear usuario charle sin permisos sudo
  • Backups automáticos en /backups

Performance

Con tu hardware (RTX 3090 + i3-10105):

  • Phi: ~50 tokens/sec
  • DeepSeek 33B: ~6-7 tokens/sec

Actualización

# Pull latest from git
git pull origin main

# Actualizar dependencias
make update

Soporte

Para errores o dudas:

  1. Ver logs: make logs
  2. Ejecutar health-check: make health-check
  3. Verificar hardware: make info

License

MIT

Autor

Setup para servidor LLM local con hardware específico.