AI News HubLIVE
站內改寫1 分鐘閱讀

待翻譯:Show HN: ComputeFence – preflight checks for rented GPU training jobs

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Notifications You must be signed in to change notification settings Fork 1 Star 0 BranchesTags Open more actions menu Latest commit History 11 Commits 11 Commits Folders and files NameName Last commit message Last commi…

來源Hacker News AI作者: Exolio_AI

AI 服務暫時不可用,以下為來源正文,待恢復後補全翻譯。

Notifications You must be signed in to change notification settings Fork 1 Star 0 BranchesTags Open more actions menu Latest commit History 11 Commits 11 Commits Folders and files NameName Last commit message Last commit date computefence computefence .gitignore .gitignore LICENSE LICENSE README.md README.md pyproject.toml pyproject.toml Repository files navigation Pre-flight validation for GPU training runs on rented infrastructure. Built specifically for RunPod, Vast.ai, Lambda, and similar providers. Install pip install computefence Usage computefence doctor computefence doctor --dataset train.csv What it checks CUDA and GPU availability HuggingFace cache volume path (catches the RunPod /root vs /workspace conflict) Accelerate GPU count vs config Dataset duplicates and missing values What it does not yet check Training script correctness Model architecture compatibility Learning rate or hyperparameter safety Runtime monitoring during the job Why this exists I burned ~£1,000 on GPU training runs that failed silently. CUDA fell back to CPU with no error. Class weights caused loss collapse. My dataset had 28,432 duplicate rows and 312 conflicting labels I only found during the rebuild. Nothing existed that caught these before the job started. So I built it. MIT license Activity Stars 0 stars Watchers 0 watching Forks 1 fork Report repository