# LLM Server - Deployment Automatizado Sistema completo de IA agentica con LLMs locales (Phi + DeepSeek 33B) en Ubuntu 22.04. ## Requisitos Mínimos - **CPU**: Intel i3-10105 (4 cores) - **RAM**: 16GB DDR4 - **GPU**: NVIDIA RTX 3090 (24GB VRAM) - **SSD**: 1TB NVMe (Samsung PM991A o similar) - **PSU**: 850W+ - **SO**: Ubuntu 22.04 LTS ## Instalación Rápida ```bash # 1. Clonar repositorio git clone https://github.com/tuuser/llm-server-setup.git cd llm-server-setup # 2. Instalar (con sudo) sudo make install # 3. Verificar estado make health-check # 4. Ver monitoreo make monitor ``` ## Comandos Disponibles ```bash make install # Instalación completa make install-quick # Sin descargar modelos make update # Actualizar código make monitor # Monitoreo en tiempo real make health-check # Verificar estado make logs # Ver logs make status # Estado de servicios make start/stop/restart # Control de servicios make backup # Hacer backup ``` ## Estructura de Directorios ``` llm-server-setup/ ├── scripts/ # Scripts de deployment │ ├── install.sh # Instalación principal │ ├── update.sh # Actualización │ ├── monitor.sh # Monitoreo │ ├── health-check.sh # Verificación │ └── backup.sh # Backup ├── config/ # Configuración │ └── server.json # Config del servidor ├── static/ # Web UI ├── workspace/ # Workspace de trabajo ├── logs/ # Archivos de log ├── backups/ # Backups automáticos ├── requirements.txt # Dependencias Python ├── Makefile # Automatización ├── .env # Variables de entorno └── README.md # Este archivo ``` ## Servicios ### Ollama (LLM Runtime) - **Puerto**: 11434 - **Modelos**: Phi (2.7B), DeepSeek Coder (33B) - **Status**: `sudo systemctl status ollama` ### LLM API (FastAPI) - **Puerto**: 8000 - **Health**: `curl http://localhost:8000/health` - **Status**: `sudo systemctl status llm-api` ## Modelos Disponibles | Modelo | Tamaño | VRAM | Velocidad | Uso | |--------|--------|------|-----------|-----| | Phi | 2.7B | 1.6GB | ⚡⚡⚡ | Rápido, simple | | DeepSeek 33B | 33B | 18GB | ⚡ | Potente, razonamiento | ## Monitoreo ```bash # Monitor en tiempo real make monitor # Ver logs make logs make logs-api # Health check make health-check ``` ## API Endpoints ```bash # Health check curl http://localhost:8000/health # Listar modelos curl http://localhost:8000/models # Generar código curl -X POST http://localhost:8000/code \ -H "Content-Type: application/json" \ -d '{"message":"Hola","model":"phi"}' ``` ## Troubleshooting ### GPU no se detecta ```bash # Verificar driver nvidia-smi # Reinstalar GPU drivers sudo make install-gpu-only ``` ### Servicios no inician ```bash # Ver logs detallados sudo journalctl -u ollama -f sudo journalctl -u llm-api -f # Reiniciar sudo systemctl restart ollama llm-api ``` ### Espacio en disco lleno ```bash # Ver uso df -h # Limpiar make clean # Hacer backup y restaurar make backup rm -rf ~/.ollama/models/* ollama pull phi:latest ``` ## Seguridad - SSH habilitado con key-based auth - Firewall: Permitir solo puertos necesarios - Usuarios: Crear usuario `charle` sin permisos sudo - Backups automáticos en `/backups` ## Performance Con tu hardware (RTX 3090 + i3-10105): - **Phi**: ~50 tokens/sec - **DeepSeek 33B**: ~6-7 tokens/sec ## Actualización ```bash # Pull latest from git git pull origin main # Actualizar dependencias make update ``` ## Soporte Para errores o dudas: 1. Ver logs: `make logs` 2. Ejecutar health-check: `make health-check` 3. Verificar hardware: `make info` ## License MIT ## Autor Setup para servidor LLM local con hardware específico.