Running Nemotron 3 Super on OpenClaw: A Free Inference Rabbit Hole

NVIDIA, OpenRouter, and Kilo are all offering free Nemotron 3 Super inference. I wired it into OpenClaw as a Discord agent, hit two silent bugs, and built a fix nobody else has published yet. Code on GitHub.

Running Nemotron 3 Super on OpenClaw: A Free Inference Rabbit Hole
Nobody said free was easy.

TL;DR: NVIDIA, OpenRouter, and Kilo are all offering free Nemotron 3 Super inference right now. I wired it into OpenClaw as a Discord agent, but OpenClaw doesn't natively support NVIDIA's API for reasoning models — so I had to adopt LiteLLM as a proxy to get it working at all. Even then, two silent bugs: the system prompt can vanish, and streaming responses come back blank. The fix is a config flag and a ~70-line Python guardrail. Nobody else has published a standalone fix — including NVIDIA's own NemoClaw project. Code and config on GitHub.

Update: A week later, more bugs. Part Deux covers config fixes, free tier limits, and what it takes to run Nemotron as a daily driver.


The pitch

Nemotron 3 Super dropped March 11. 120 billion parameters, 12 billion active per forward pass, and — here's the part that got me — free inference through multiple providers. NVIDIA NIM gives you 40 requests per minute on their Developer Program. OpenRouter has it on their free tier. Kilo hands you $5 in credits on signup. No end dates published on any of them.

A free 120B reasoning model is hard to ignore. I wanted to run it as my Discord agent for a while and see how it holds up. Three days later I had a working fix and a mass of Docker logs I'll never unsee.

For what it's worth, the model is solid. The problems I ran into have nothing to do with Nemotron itself — they're all in the plumbing between it and OpenClaw.

The stack

My agent "Cass" runs on OpenClaw, talks through Discord, and normally runs on Anthropic models. OpenClaw doesn't natively support NVIDIA NIM for reasoning models, so I needed a proxy. LiteLLM does this — it sits between OpenClaw and whatever backend you point it at, speaks OpenAI-compatible on both sides, and handles the translation. I'm new to it, but it earned its keep on this project. Model routing, Prometheus metrics, spend tracking, all out of the box.

Discord ← OpenClaw Gateway ← LiteLLM Proxy (:4000) ← NVIDIA NIM / OpenRouter / LM Studio

I wired up four backends that I can swap from Discord with /model nemotron-nvidia or /model nemotron-local. LiteLLM runs as Docker alongside Postgres and Prometheus.

One important note: I'm a systems engineer, not a Python developer. I think in architecture diagrams, not async generators. So I fired up Claude Opus and spent three sessions debugging this together. Opus wrote the code. I made the architectural decisions and did the "wait, verify that claim before we ship it" work that keeps AI-assisted development from going off the rails.

What broke

Two things. Both silent. No errors, no warnings, no indication that anything was wrong — just an agent that sometimes responded with nothing.

The system prompt problem

This one came from research, not firsthand debugging. Rogério Richa wrote a great Medium post about the exact same issue with Ollama. When OpenClaw sees reasoning: true on a model, it sends the system prompt as role: "developer". That's an OpenAI-specific role. Providers that don't recognize it just drop the message. Silently. Your agent's entire personality — every instruction, every behavioral rule — gone without a trace.

I couldn't confirm this was happening in my LiteLLM→NIM path specifically, but the fix is a single config flag with zero downside:

{
  "reasoning": true,
  "compat": {
    "supportsDeveloperRole": false,
    "supportsReasoningEffort": true
  }
}

That forces role: "system", which every provider understands. Applied it and moved on.

The blank response problem

This one I could confirm, and it's the reason this post exists.

When Nemotron streams a response through NVIDIA NIM, the thinking phase sends identical text in both reasoning_content and content. The same tokens, duplicated into both fields. When the actual answer starts, it only shows up in content. OpenClaw sees this overlapping mess and either renders blank messages or leaks raw reasoning into the chat.

I spent a while thinking I was misconfiguring something. I wasn't. NVIDIA's own NemoClaw integration has the same bug, filed two days before I found it. OpenClaw #27806 describes the same thing. So does #14071. Nobody has published a fix.

The fix

LiteLLM has a guardrail system — the same hook mechanism that enterprise security products like Palo Alto Networks AIRS use to inspect streaming responses. I used it for something less glamorous: deduplicating fields.

The guardrail intercepts every SSE chunk in the stream. The logic is almost embarrassingly simple:

  • If reasoning and content contain the same text, suppress content. That's the thinking phase.
  • The moment content diverges from reasoning, flip: suppress reasoning, pass only content. That's the answer.
class ReasoningGuardrail(CustomGuardrail):
    async def async_post_call_streaming_iterator_hook(
        self, user_api_key_dict, response, request_data
    ):
        saw_answer = False
        async for chunk in response:
            delta = chunk.choices[0].delta
            reasoning = getattr(delta, "reasoning", None)
            content = getattr(delta, "content", None)

            if reasoning and content and reasoning == content and not saw_answer:
                mod = copy.deepcopy(chunk)
                mod.choices[0].delta.content = None
                mod.choices[0].delta.reasoning_content = reasoning
                yield mod
                continue

            if content and (not reasoning or reasoning != content):
                saw_answer = True
                mod = copy.deepcopy(chunk)
                mod.choices[0].delta.reasoning_content = None
                yield mod
                continue

            yield chunk

Register it in the LiteLLM config, volume-mount the Python file into the Docker container, restart. Thinking goes to the reasoning panel, answers go to chat. Done.

guardrails:
  - guardrail_name: "reasoning-fixer"
    litellm_params:
      guardrail: reasoning_guardrail.ReasoningGuardrail
      mode: "post_call"
      default_on: true

The detour I could have skipped

We also spent a few hours building a pre-call hook to translate OpenClaw's /think commands into NVIDIA's native chat_template_kwargs format. It involved debugging an import error where UserAPIKeyAuth moved between LiteLLM versions, discovering that CustomLogger callbacks don't fire for streaming (LiteLLM #9639), re-registering everything as a CustomGuardrail, and finally getting it to load cleanly.

Then we checked the LiteLLM request logs and found it already translates reasoning_effort into chat_template_kwargs natively for nvidia_nim/ models. The hook was redundant before we finished writing it. That's how it goes sometimes.

What works

From Discord, after all of this:

  • /think level low — reasoning in the collapsible panel, clean answer in chat
  • /think level high — extended reasoning, same separation
  • /think level off — disables thinking entirely, answer goes straight to content. The model does occasionally narrate its reasoning in plain text when it doesn't have a thinking block. That's a Nemotron personality trait, not a proxy bug.
  • /model switching — swap between NVIDIA NIM, OpenRouter, LM Studio, Kilo from Discord
  • Tool calls — work through LiteLLM across all backends
  • System prompt — delivers correctly with compat flags

Before you implement this

This is a bridge. The upstream issues (OpenClaw #27806, NemoClaw #247) are open. When OpenClaw handles reasoning_content natively, this guardrail becomes dead code. Check those issues first.

The guardrail runs on everything. default_on: true applies it to all models through LiteLLM. It's harmless for non-reasoning models, but a production deployment should scope it.

The compat flags are semi-documented. They're in OpenClaw's source and used across GitHub issues, but not in the official docs. They could change.

The bigger picture

Nemotron 3 Super launched March 11. NemoClaw launched at GTC on March 16. The bug report was filed March 17. This fix was working March 19.

The whole ecosystem — OpenClaw, NemoClaw, LiteLLM, the community — is still working out how reasoning models should behave in the OpenAI-compatible streaming format. We're doing at the proxy layer what NVIDIA's OpenShell gateway does at the kernel level. Same pattern, ~70 lines of Python, and a mass of Docker logs.

Code and config: github.com/cassthebandit/nemotron-openclaw-fix

Update: This fix was the starting point. Part Deux covers what happened next — config fixes, reasoning budget research, free tier provider limits, and the LM Studio vs. LiteLLM comparison.


Built with Claude. Directed by me.


Daniel Soteldo is COO and Co-Founder of Revelus Dermatology in Austin, TX.