187 lines
3.9 KiB
Markdown
187 lines
3.9 KiB
Markdown
# LLM Server - Deployment Automatizado
|
|
|
|
Sistema completo de IA agentica con LLMs locales (Phi + DeepSeek 33B) en Ubuntu 22.04.
|
|
|
|
## Requisitos Mínimos
|
|
|
|
- **CPU**: Intel i3-10105 (4 cores)
|
|
- **RAM**: 16GB DDR4
|
|
- **GPU**: NVIDIA RTX 3090 (24GB VRAM)
|
|
- **SSD**: 1TB NVMe (Samsung PM991A o similar)
|
|
- **PSU**: 850W+
|
|
- **SO**: Ubuntu 22.04 LTS
|
|
|
|
## Instalación Rápida
|
|
|
|
```bash
|
|
# 1. Clonar repositorio
|
|
git clone https://github.com/tuuser/llm-server-setup.git
|
|
cd llm-server-setup
|
|
|
|
# 2. Instalar (con sudo)
|
|
sudo make install
|
|
|
|
# 3. Verificar estado
|
|
make health-check
|
|
|
|
# 4. Ver monitoreo
|
|
make monitor
|
|
```
|
|
|
|
## Comandos Disponibles
|
|
|
|
```bash
|
|
make install # Instalación completa
|
|
make install-quick # Sin descargar modelos
|
|
make update # Actualizar código
|
|
make monitor # Monitoreo en tiempo real
|
|
make health-check # Verificar estado
|
|
make logs # Ver logs
|
|
make status # Estado de servicios
|
|
make start/stop/restart # Control de servicios
|
|
make backup # Hacer backup
|
|
```
|
|
|
|
## Estructura de Directorios
|
|
|
|
```
|
|
llm-server-setup/
|
|
├── scripts/ # Scripts de deployment
|
|
│ ├── install.sh # Instalación principal
|
|
│ ├── update.sh # Actualización
|
|
│ ├── monitor.sh # Monitoreo
|
|
│ ├── health-check.sh # Verificación
|
|
│ └── backup.sh # Backup
|
|
├── config/ # Configuración
|
|
│ └── server.json # Config del servidor
|
|
├── static/ # Web UI
|
|
├── workspace/ # Workspace de trabajo
|
|
├── logs/ # Archivos de log
|
|
├── backups/ # Backups automáticos
|
|
├── requirements.txt # Dependencias Python
|
|
├── Makefile # Automatización
|
|
├── .env # Variables de entorno
|
|
└── README.md # Este archivo
|
|
```
|
|
|
|
## Servicios
|
|
|
|
### Ollama (LLM Runtime)
|
|
- **Puerto**: 11434
|
|
- **Modelos**: Phi (2.7B), DeepSeek Coder (33B)
|
|
- **Status**: `sudo systemctl status ollama`
|
|
|
|
### LLM API (FastAPI)
|
|
- **Puerto**: 8000
|
|
- **Health**: `curl http://localhost:8000/health`
|
|
- **Status**: `sudo systemctl status llm-api`
|
|
|
|
## Modelos Disponibles
|
|
|
|
| Modelo | Tamaño | VRAM | Velocidad | Uso |
|
|
|--------|--------|------|-----------|-----|
|
|
| Phi | 2.7B | 1.6GB | ⚡⚡⚡ | Rápido, simple |
|
|
| DeepSeek 33B | 33B | 18GB | ⚡ | Potente, razonamiento |
|
|
|
|
## Monitoreo
|
|
|
|
```bash
|
|
# Monitor en tiempo real
|
|
make monitor
|
|
|
|
# Ver logs
|
|
make logs
|
|
make logs-api
|
|
|
|
# Health check
|
|
make health-check
|
|
```
|
|
|
|
## API Endpoints
|
|
|
|
```bash
|
|
# Health check
|
|
curl http://localhost:8000/health
|
|
|
|
# Listar modelos
|
|
curl http://localhost:8000/models
|
|
|
|
# Generar código
|
|
curl -X POST http://localhost:8000/code \
|
|
-H "Content-Type: application/json" \
|
|
-d '{"message":"Hola","model":"phi"}'
|
|
```
|
|
|
|
## Troubleshooting
|
|
|
|
### GPU no se detecta
|
|
```bash
|
|
# Verificar driver
|
|
nvidia-smi
|
|
|
|
# Reinstalar GPU drivers
|
|
sudo make install-gpu-only
|
|
```
|
|
|
|
### Servicios no inician
|
|
```bash
|
|
# Ver logs detallados
|
|
sudo journalctl -u ollama -f
|
|
sudo journalctl -u llm-api -f
|
|
|
|
# Reiniciar
|
|
sudo systemctl restart ollama llm-api
|
|
```
|
|
|
|
### Espacio en disco lleno
|
|
```bash
|
|
# Ver uso
|
|
df -h
|
|
|
|
# Limpiar
|
|
make clean
|
|
|
|
# Hacer backup y restaurar
|
|
make backup
|
|
rm -rf ~/.ollama/models/*
|
|
ollama pull phi:latest
|
|
```
|
|
|
|
## Seguridad
|
|
|
|
- SSH habilitado con key-based auth
|
|
- Firewall: Permitir solo puertos necesarios
|
|
- Usuarios: Crear usuario `charle` sin permisos sudo
|
|
- Backups automáticos en `/backups`
|
|
|
|
## Performance
|
|
|
|
Con tu hardware (RTX 3090 + i3-10105):
|
|
|
|
- **Phi**: ~50 tokens/sec
|
|
- **DeepSeek 33B**: ~6-7 tokens/sec
|
|
|
|
## Actualización
|
|
|
|
```bash
|
|
# Pull latest from git
|
|
git pull origin main
|
|
|
|
# Actualizar dependencias
|
|
make update
|
|
```
|
|
|
|
## Soporte
|
|
|
|
Para errores o dudas:
|
|
1. Ver logs: `make logs`
|
|
2. Ejecutar health-check: `make health-check`
|
|
3. Verificar hardware: `make info`
|
|
|
|
## License
|
|
|
|
MIT
|
|
|
|
## Autor
|
|
|
|
Setup para servidor LLM local con hardware específico. |