How does ATH train multimodal models without visual hallucination?
We utilize Direct Preference Optimization adapted for Vision (DPO-V) coupled with pixel-grounded bounding-box loss. Our synthetic curation pipeline automatically injects negative examples with modified visual components, teaching the model to refuse to identify non-existent items with mathematical consistency.
What open-weights architectures can be customized?
ATH fine-tunes and aligns Llama 3.2 Vision (11B & 90B), Qwen2-VL, Pixtral 12B, DeepSeek-VL, and custom ViT-SigLIP adapter backbones. We also build custom lightweight Vision-Language adapters for proprietary on-premise foundation LLMs.
How do you handle private company data and HIPAA/SOC 2 compliance?
All training is executed in single-tenant, customer-dedicated VPCs or air-gapped on-premise GPU clusters. ATH never aggregates data across customers, enforces strict cryptographic Zero Data Retention, and provides audited SOC 2 Type II and HIPAA attestation reports.
Can ATH models run on on-premise GPU clusters and air-gapped VPCs?
Yes. We containerize models using Docker and Kubernetes (KServe, vLLM, Triton) for seamless deployment onto NVIDIA H100/A100 clusters, AWS GovCloud, or on-premise high-density compute nodes without requiring external internet calls.
What is the typical timeline and GPU requirement for training?
A standard domain-specific multimodal adaptation takes 4 to 8 weeks, including dataset synthesis, cross-attention alignment, and DPO-V tuning. GPU allocations range from an 8x H100 node for LoRA/QLoRA adapter fine-tuning to dedicated 32x-64x H100 clusters for full-parameter multimodal pre-training.