Running Nemotron 3 Super on OpenClaw: A Free Inference Rabbit Hole
NVIDIA, OpenRouter, and Kilo are all offering free Nemotron 3 Super inference. I wired it into OpenClaw as a Discord agent, hit two silent bugs, and built a fix nobody else has published yet. Code on GitHub.
TL;DR: NVIDIA, OpenRouter, and Kilo are all offering free Nemotron 3 Super inference right now. I wired it into OpenClaw as a Discord agent, but OpenClaw doesn't natively support NVIDIA's API for reasoning models — so I had to adopt LiteLLM as a proxy to get it working at all. Even then, two silent bugs: the system prompt can vanish, and streaming responses come back blank. The fix is a config flag and a ~70-line Python guardrail. Nobody else has published a standalone fix — including NVIDIA's own NemoClaw project. Code and config on GitHub.
Update: A week later, more bugs. Part Deux covers config fixes, free tier limits, and what it takes to run Nemotron as a daily driver.
The pitch
Nemotron 3 Super dropped March 11. 120 billion parameters, 12 billion active per forward pass, and — here's the part that got me — free inference through multiple providers. NVIDIA NIM gives you 40 requests per minute on their Developer Program. OpenRouter has it on their free tier. Kilo hands you $5 in credits on signup. No end dates published on any of them.
A free 120B reasoning model is hard to ignore. I wanted to run it as my Discord agent for a while and see how it holds up. Three days later I had a working fix and a mass of Docker logs I'll never unsee.
For what it's worth, the model is solid. The problems I ran into have nothing to do with Nemotron itself — they're all in the plumbing between it and OpenClaw.
The stack
My agent "Cass" runs on OpenClaw, talks through Discord, and normally runs on Anthropic models. OpenClaw doesn't natively support NVIDIA NIM for reasoning models, so I needed a proxy. LiteLLM does this — it sits between OpenClaw and whatever backend you point it at, speaks OpenAI-compatible on both sides, and handles the translation. I'm new to it, but it earned its keep on this project. Model routing, Prometheus metrics, spend tracking, all out of the box.
Discord ← OpenClaw Gateway ← LiteLLM Proxy (:4000) ← NVIDIA NIM / OpenRouter / LM Studio
I wired up four backends that I can swap from Discord with /model nemotron-nvidia or /model nemotron-local. LiteLLM runs as Docker alongside Postgres and Prometheus.
One important note: I'm a systems engineer, not a Python developer. I think in architecture diagrams, not async generators. So I fired up Claude Opus and spent three sessions debugging this together. Opus wrote the code. I made the architectural decisions and did the "wait, verify that claim before we ship it" work that keeps AI-assisted development from going off the rails.
What broke
Two things. Both silent. No errors, no warnings, no indication that anything was wrong — just an agent that sometimes responded with nothing.
The system prompt problem
This one came from research, not firsthand debugging. Rogério Richa wrote a great Medium post about the exact same issue with Ollama. When OpenClaw sees reasoning: true on a model, it sends the system prompt as role: "developer". That's an OpenAI-specific role. Providers that don't recognize it just drop the message. Silently. Your agent's entire personality — every instruction, every behavioral rule — gone without a trace.
I couldn't confirm this was happening in my LiteLLM→NIM path specifically, but the fix is a single config flag with zero downside:
{
"reasoning": true,
"compat": {
"supportsDeveloperRole": false,
"supportsReasoningEffort": true
}
}
That forces role: "system", which every provider understands. Applied it and moved on.
The blank response problem
This one I could confirm, and it's the reason this post exists.
When Nemotron streams a response through NVIDIA NIM, the thinking phase sends identical text in both reasoning_content and content. The same tokens, duplicated into both fields. When the actual answer starts, it only shows up in content. OpenClaw sees this overlapping mess and either renders blank messages or leaks raw reasoning into the chat.
I spent a while thinking I was misconfiguring something. I wasn't. NVIDIA's own NemoClaw integration has the same bug, filed two days before I found it. OpenClaw #27806 describes the same thing. So does #14071. Nobody has published a fix.
The fix
LiteLLM has a guardrail system — the same hook mechanism that enterprise security products like Palo Alto Networks AIRS use to inspect streaming responses. I used it for something less glamorous: deduplicating fields.
The guardrail intercepts every SSE chunk in the stream. The logic is almost embarrassingly simple:
- If
reasoningandcontentcontain the same text, suppresscontent. That's the thinking phase. - The moment
contentdiverges fromreasoning, flip: suppress reasoning, pass onlycontent. That's the answer.
class ReasoningGuardrail(CustomGuardrail):
async def async_post_call_streaming_iterator_hook(
self, user_api_key_dict, response, request_data
):
saw_answer = False
async for chunk in response:
delta = chunk.choices[0].delta
reasoning = getattr(delta, "reasoning", None)
content = getattr(delta, "content", None)
if reasoning and content and reasoning == content and not saw_answer:
mod = copy.deepcopy(chunk)
mod.choices[0].delta.content = None
mod.choices[0].delta.reasoning_content = reasoning
yield mod
continue
if content and (not reasoning or reasoning != content):
saw_answer = True
mod = copy.deepcopy(chunk)
mod.choices[0].delta.reasoning_content = None
yield mod
continue
yield chunk
Register it in the LiteLLM config, volume-mount the Python file into the Docker container, restart. Thinking goes to the reasoning panel, answers go to chat. Done.
guardrails:
- guardrail_name: "reasoning-fixer"
litellm_params:
guardrail: reasoning_guardrail.ReasoningGuardrail
mode: "post_call"
default_on: true
The detour I could have skipped
We also spent a few hours building a pre-call hook to translate OpenClaw's /think commands into NVIDIA's native chat_template_kwargs format. It involved debugging an import error where UserAPIKeyAuth moved between LiteLLM versions, discovering that CustomLogger callbacks don't fire for streaming (LiteLLM #9639), re-registering everything as a CustomGuardrail, and finally getting it to load cleanly.
Then we checked the LiteLLM request logs and found it already translates reasoning_effort into chat_template_kwargs natively for nvidia_nim/ models. The hook was redundant before we finished writing it. That's how it goes sometimes.
What works
From Discord, after all of this:
/think level low— reasoning in the collapsible panel, clean answer in chat/think level high— extended reasoning, same separation/think level off— disables thinking entirely, answer goes straight to content. The model does occasionally narrate its reasoning in plain text when it doesn't have a thinking block. That's a Nemotron personality trait, not a proxy bug./modelswitching — swap between NVIDIA NIM, OpenRouter, LM Studio, Kilo from Discord- Tool calls — work through LiteLLM across all backends
- System prompt — delivers correctly with compat flags
Before you implement this
This is a bridge. The upstream issues (OpenClaw #27806, NemoClaw #247) are open. When OpenClaw handles reasoning_content natively, this guardrail becomes dead code. Check those issues first.
The guardrail runs on everything. default_on: true applies it to all models through LiteLLM. It's harmless for non-reasoning models, but a production deployment should scope it.
The compat flags are semi-documented. They're in OpenClaw's source and used across GitHub issues, but not in the official docs. They could change.
The bigger picture
Nemotron 3 Super launched March 11. NemoClaw launched at GTC on March 16. The bug report was filed March 17. This fix was working March 19.
The whole ecosystem — OpenClaw, NemoClaw, LiteLLM, the community — is still working out how reasoning models should behave in the OpenAI-compatible streaming format. We're doing at the proxy layer what NVIDIA's OpenShell gateway does at the kernel level. Same pattern, ~70 lines of Python, and a mass of Docker logs.
Code and config: github.com/cassthebandit/nemotron-openclaw-fix
Update: This fix was the starting point. Part Deux covers what happened next — config fixes, reasoning budget research, free tier provider limits, and the LM Studio vs. LiteLLM comparison.
Built with Claude. Directed by me.
Daniel Soteldo is COO and Co-Founder of Revelus Dermatology in Austin, TX.