Files
llm-server-setup/README.md
T
2026-09-20 13:43:58 -03:00

187 lines
3.9 KiB
Markdown

# LLM Server - Deployment Automatizado
Sistema completo de IA agentica con LLMs locales (Phi + DeepSeek 33B) en Ubuntu 22.04.
## Requisitos Mínimos
- **CPU**: Intel i3-10105 (4 cores)
- **RAM**: 16GB DDR4
- **GPU**: NVIDIA RTX 3090 (24GB VRAM)
- **SSD**: 1TB NVMe (Samsung PM991A o similar)
- **PSU**: 850W+
- **SO**: Ubuntu 22.04 LTS
## Instalación Rápida
```bash
# 1. Clonar repositorio
git clone https://github.com/tuuser/llm-server-setup.git
cd llm-server-setup
# 2. Instalar (con sudo)
sudo make install
# 3. Verificar estado
make health-check
# 4. Ver monitoreo
make monitor
```
## Comandos Disponibles
```bash
make install # Instalación completa
make install-quick # Sin descargar modelos
make update # Actualizar código
make monitor # Monitoreo en tiempo real
make health-check # Verificar estado
make logs # Ver logs
make status # Estado de servicios
make start/stop/restart # Control de servicios
make backup # Hacer backup
```
## Estructura de Directorios
```
llm-server-setup/
├── scripts/ # Scripts de deployment
│ ├── install.sh # Instalación principal
│ ├── update.sh # Actualización
│ ├── monitor.sh # Monitoreo
│ ├── health-check.sh # Verificación
│ └── backup.sh # Backup
├── config/ # Configuración
│ └── server.json # Config del servidor
├── static/ # Web UI
├── workspace/ # Workspace de trabajo
├── logs/ # Archivos de log
├── backups/ # Backups automáticos
├── requirements.txt # Dependencias Python
├── Makefile # Automatización
├── .env # Variables de entorno
└── README.md # Este archivo
```
## Servicios
### Ollama (LLM Runtime)
- **Puerto**: 11434
- **Modelos**: Phi (2.7B), DeepSeek Coder (33B)
- **Status**: `sudo systemctl status ollama`
### LLM API (FastAPI)
- **Puerto**: 8000
- **Health**: `curl http://localhost:8000/health`
- **Status**: `sudo systemctl status llm-api`
## Modelos Disponibles
| Modelo | Tamaño | VRAM | Velocidad | Uso |
|--------|--------|------|-----------|-----|
| Phi | 2.7B | 1.6GB | ⚡⚡⚡ | Rápido, simple |
| DeepSeek 33B | 33B | 18GB | ⚡ | Potente, razonamiento |
## Monitoreo
```bash
# Monitor en tiempo real
make monitor
# Ver logs
make logs
make logs-api
# Health check
make health-check
```
## API Endpoints
```bash
# Health check
curl http://localhost:8000/health
# Listar modelos
curl http://localhost:8000/models
# Generar código
curl -X POST http://localhost:8000/code \
-H "Content-Type: application/json" \
-d '{"message":"Hola","model":"phi"}'
```
## Troubleshooting
### GPU no se detecta
```bash
# Verificar driver
nvidia-smi
# Reinstalar GPU drivers
sudo make install-gpu-only
```
### Servicios no inician
```bash
# Ver logs detallados
sudo journalctl -u ollama -f
sudo journalctl -u llm-api -f
# Reiniciar
sudo systemctl restart ollama llm-api
```
### Espacio en disco lleno
```bash
# Ver uso
df -h
# Limpiar
make clean
# Hacer backup y restaurar
make backup
rm -rf ~/.ollama/models/*
ollama pull phi:latest
```
## Seguridad
- SSH habilitado con key-based auth
- Firewall: Permitir solo puertos necesarios
- Usuarios: Crear usuario `charle` sin permisos sudo
- Backups automáticos en `/backups`
## Performance
Con tu hardware (RTX 3090 + i3-10105):
- **Phi**: ~50 tokens/sec
- **DeepSeek 33B**: ~6-7 tokens/sec
## Actualización
```bash
# Pull latest from git
git pull origin main
# Actualizar dependencias
make update
```
## Soporte
Para errores o dudas:
1. Ver logs: `make logs`
2. Ejecutar health-check: `make health-check`
3. Verificar hardware: `make info`
## License
MIT
## Autor
Setup para servidor LLM local con hardware específico.