Saptiva
Client

Catálogo de modelos

Los modelos disponibles en tu infraestructura y su almacenamiento. Los recursos en vivo (VRAM) se gestionan en Modelos.

Almacenamiento usado
231 GB
11% de 2048 GB · NVMe RAID (2 TB)
Modelos descargados
9
1817 GB libres
Activos (no borrables)
4
En uso por modelos de la organización
Actualización disponible
3
Ver abajo
ModeloProveedorTipoParamsTamañoVRAM est.Estado
Llama 3.1 8B Instruct
MetaInstrucción8B16.1 GB18 GB
Activo
Qwen 2.5 7B Instruct
QwenInstrucción7B14.7 GB16 GB
Descargado
Mistral 7B Instruct v0.3
Mistral AIInstrucción7B14.5 GB15 GB
Descargado
Llama 3.1 70B Instruct
MetaInstrucción70B140.0 GB160 GB
Descargando 67%
Pixtral 12B
Mistral AIVisión12B24.5 GB28 GB
Activo
Qwen 2 VL 7B
QwenVisión7B15.8 GB18 GB
Descargado
Nougat Base (OCR)
MetaOCR0.35B1.4 GB4 GB
Activo
BGE M3 (Embeddings)
NVIDIAEmbeddings0.57B2.3 GB3 GB
Activo
BGE Large EN v1.5
NVIDIAEmbeddings0.34B1.3 GB2 GB
Descargado
meta-llama/Llama-3.2-3B-Instruct
MetaInstrucción3B6.2 GB8 GB
meta-llama/Llama-3.3-70B-Instruct
MetaInstrucción70B142.0 GB160 GB
Qwen/Qwen2.5-14B-Instruct
QwenInstrucción14B28.4 GB32 GB
Qwen/Qwen2.5-32B-Instruct
QwenInstrucción32B65.0 GB72 GB
Qwen/Qwen2.5-Coder-32B-Instruct
QwenInstrucción32B65.0 GB72 GB
google/gemma-2-9b-it
GoogleInstrucción9B18.2 GB22 GB
google/gemma-2-27b-it
GoogleInstrucción27B54.5 GB62 GB
google/gemma-3-12b-it
GoogleVisión12B24.6 GB28 GB
google/paligemma-3b-mix-448
GoogleVisión3B6.1 GB8 GB
nvidia/NV-Embed-v2
NVIDIAEmbeddings7.85B31.4 GB36 GB
nvidia/Llama-3.1-Nemotron-70B-Instruct
NVIDIAInstrucción70B140.0 GB160 GB
nvidia/Nemotron-Mini-4B-Instruct
NVIDIAInstrucción4B8.1 GB10 GB
21 modelos en el catálogo · 9 ya en tu infra

Actualizaciones disponibles

· 3
  • Qwen 2.5 7B Instruct
    Qwen2.5-7B-Instruct-1MSoporta 1M tokens de contexto
  • Mistral 7B Instruct v0.3
    Mistral-7B-Instruct-v0.4Mejor seguimiento de instrucciones
  • BGE Large EN v1.5
    bge-large-en-v1.5 (rev 2026-04)Recalibrado, mejor MTEB