Gerald Stanje
Gerald001
·
AI & ML interests
None yet
Recent Activity
new activity 3 days ago
openai/gpt-oss-safeguard-20b:release the BF16 weights or nvfp4 new activity 4 days ago
openai/gpt-oss-20b:MXFP4 utilization over NVFP4 new activity 4 days ago
Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice:Inference latencyOrganizations
release the BF16 weights or nvfp4
#7 opened 3 days ago
by
Gerald001
MXFP4 utilization over NVFP4
👍 2
2
#121 opened 11 months ago
by
pprovins
Inference latency
1
#24 opened 6 months ago
by
popkek00
Question about inference speed
2
#10 opened 11 months ago
by
cX1y
I want transcription not translation
👍 5
2
#29 opened 7 months ago
by
smartire
What's the best way to run this in production for online serving?
1
#45 opened 27 days ago
by
fikrikarim
Any plans for gpt oss 20b?
2
#1 opened 9 months ago
by
andhakanoon
Upload TRT model for Nvidia H100
#5 opened 4 months ago
by
robustdev
Upload TRT model for Nvidia H100
#4 opened 4 months ago
by
robustdev
Upload TRT model for Nvidia H200
#3 opened 4 months ago
by
robustdev
Upload TRT model for Nvidia L40S
#2 opened 4 months ago
by
robustdev
Upload ONNX model for Nvidia L40S
#1 opened 4 months ago
by
robustdev
Eagle3 for 20b
🚀 1
3
#4 opened 9 months ago
by
cr-boostrun
how to disable the reasoning mode?
👍 13
11
#50 opened 12 months ago
by
szzzzz
GGUF is very slow for some reason
2
#12 opened 11 months ago
by
ineersa
Model conversion info
11
#9 opened 5 months ago
by
Gerald001
NVIDIA L40S GPU's for MXFP4 quantization
6
#100 opened 12 months ago
by
lordim
How to turn off thinking mode
👍🔥 7
15
#86 opened 12 months ago
by
Gierry
REASONING SETTING GUIDE 📚
😔👍 4
28
#28 opened 12 months ago
by
xbruce22