> ## Content Index
> Fetch the complete content index at: https://infer.blog/llms.txt
> Use this file to discover other available public pages before exploring further.

# Running Nemotron 3 Super on OpenClaw: A Free Inference Rabbit Hole
- URL: https://infer.blog/running-nemotron-3-super-on-openclaw-a-free-inference-rabbit-hole/
- Published: 2026-03-19T23:04:34.000Z
- Updated: 2026-03-27T00:41:40.000Z
- Description: NVIDIA, OpenRouter, and Kilo are all offering free Nemotron 3 Super inference. I wired it into OpenClaw as a Discord agent, hit two silent bugs, and built a fix nobody else has published yet. Code on GitHub.
- Author: Daniel Soteldo
- Tags: AI, LiteLLM, Nemotron, OpenClaw, Nvidia NIM

**TL;DR:** NVIDIA, OpenRouter, and Kilo are all offering free Nemotron 3 Super inference right now. I wired it into OpenClaw as a Discord agent, but OpenClaw doesn't natively support NVIDIA's API for reasoning models — so I had to adopt [LiteLLM](https://github.com/BerriAI/litellm?ref=infer.blog) as a proxy to get it working at all. Even then, two silent bugs: the system prompt can vanish, and streaming responses come back blank. The fix is a config flag and a \~70-line Python guardrail. Nobody else has published a standalone fix — including [NVIDIA's own NemoClaw project](https://github.com/NVIDIA/NemoClaw/issues/247?ref=infer.blog). Code and config [on GitHub](https://github.com/cassthebandit/nemotron-openclaw-fix?ref=infer.blog).

**Update:** A week later, more bugs. [Part Deux](https://infer.blog/running-nemotron-3-super-on-openclaw-part-deux/) covers config fixes, free tier limits, and what it takes to run Nemotron as a daily driver.

---

## The pitch

[Nemotron 3 Super](https://research.nvidia.com/labs/nemotron/Nemotron-3-Super/?ref=infer.blog) dropped March 11\. 120 billion parameters, 12 billion active per forward pass, and — here's the part that got me — free inference through multiple providers. [NVIDIA NIM](https://build.nvidia.com/?ref=infer.blog) gives you 40 requests per minute on their Developer Program. [OpenRouter](https://openrouter.ai/nvidia/nemotron-3-super-120b-a12b:free?ref=infer.blog) has it on their free tier. [Kilo](https://kilo.ai/gateway?ref=infer.blog) hands you $5 in credits on signup. No end dates published on any of them.

A free 120B reasoning model is hard to ignore. I wanted to run it as my Discord agent for a while and see how it holds up. Three days later I had a working fix and a mass of Docker logs I'll never unsee.

For what it's worth, the model is solid. The problems I ran into have nothing to do with Nemotron itself — they're all in the plumbing between it and OpenClaw.

## The stack

My agent "Cass" runs on [OpenClaw](https://github.com/openclaw/openclaw?ref=infer.blog), talks through Discord, and normally runs on Anthropic models. OpenClaw doesn't natively support NVIDIA NIM for reasoning models, so I needed a proxy. [LiteLLM](https://github.com/BerriAI/litellm?ref=infer.blog) does this — it sits between OpenClaw and whatever backend you point it at, speaks OpenAI-compatible on both sides, and handles the translation. I'm new to it, but it earned its keep on this project. Model routing, Prometheus metrics, spend tracking, all out of the box.

```
Discord ← OpenClaw Gateway ← LiteLLM Proxy (:4000) ← NVIDIA NIM / OpenRouter / LM Studio

```

I wired up four backends that I can swap from Discord with `/model nemotron-nvidia` or `/model nemotron-local`. LiteLLM runs as Docker alongside Postgres and Prometheus.

One important note: I'm a systems engineer, not a Python developer. I think in architecture diagrams, not async generators. So I fired up Claude Opus and spent three sessions debugging this together. Opus wrote the code. I made the architectural decisions and did the "wait, verify that claim before we ship it" work that keeps AI-assisted development from going off the rails.

## What broke

Two things. Both silent. No errors, no warnings, no indication that anything was wrong — just an agent that sometimes responded with nothing.

### The system prompt problem

This one came from research, not firsthand debugging. [Rogério Richa wrote a great Medium post](https://medium.com/@rogerio.a.r/setting-up-a-private-local-llm-with-ollama-for-use-with-openclaw-a-tale-of-silent-failures-01cadfee717f?ref=infer.blog) about the exact same issue with Ollama. When OpenClaw sees `reasoning: true` on a model, it sends the system prompt as `role: "developer"`. That's an OpenAI-specific role. Providers that don't recognize it just drop the message. Silently. Your agent's entire personality — every instruction, every behavioral rule — gone without a trace.

I couldn't confirm this was happening in my LiteLLM→NIM path specifically, but the fix is a single config flag with zero downside:

```json
{
  "reasoning": true,
  "compat": {
    "supportsDeveloperRole": false,
    "supportsReasoningEffort": true
  }
}

```

That forces `role: "system"`, which every provider understands. Applied it and moved on.

### The blank response problem

This one I could confirm, and it's the reason this post exists.

When Nemotron streams a response through NVIDIA NIM, the thinking phase sends identical text in both `reasoning_content` and `content`. The same tokens, duplicated into both fields. When the actual answer starts, it only shows up in `content`. OpenClaw sees this overlapping mess and either renders blank messages or leaks raw reasoning into the chat.

I spent a while thinking I was misconfiguring something. I wasn't. [NVIDIA's own NemoClaw integration has the same bug](https://github.com/NVIDIA/NemoClaw/issues/247?ref=infer.blog), filed two days before I found it. [OpenClaw #27806](https://github.com/openclaw/openclaw/issues/27806?ref=infer.blog) describes the same thing. [So does #14071](https://github.com/openclaw/openclaw/issues/14071?ref=infer.blog). Nobody has published a fix.

## The fix

LiteLLM has a [guardrail system](https://docs.litellm.ai/docs/proxy/guardrails/custom%5Fguardrail?ref=infer.blog) — the same hook mechanism that enterprise security products like Palo Alto Networks AIRS use to inspect streaming responses. I used it for something less glamorous: deduplicating fields.

The guardrail intercepts every SSE chunk in the stream. The logic is almost embarrassingly simple:

- If `reasoning` and `content` contain the same text, suppress `content`. That's the thinking phase.
- The moment `content` diverges from `reasoning`, flip: suppress reasoning, pass only `content`. That's the answer.

```python
class ReasoningGuardrail(CustomGuardrail):
    async def async_post_call_streaming_iterator_hook(
        self, user_api_key_dict, response, request_data
    ):
        saw_answer = False
        async for chunk in response:
            delta = chunk.choices[0].delta
            reasoning = getattr(delta, "reasoning", None)
            content = getattr(delta, "content", None)

            if reasoning and content and reasoning == content and not saw_answer:
                mod = copy.deepcopy(chunk)
                mod.choices[0].delta.content = None
                mod.choices[0].delta.reasoning_content = reasoning
                yield mod
                continue

            if content and (not reasoning or reasoning != content):
                saw_answer = True
                mod = copy.deepcopy(chunk)
                mod.choices[0].delta.reasoning_content = None
                yield mod
                continue

            yield chunk

```

Register it in the LiteLLM config, volume-mount the Python file into the Docker container, restart. Thinking goes to the reasoning panel, answers go to chat. Done.

```yaml
guardrails:
  - guardrail_name: "reasoning-fixer"
    litellm_params:
      guardrail: reasoning_guardrail.ReasoningGuardrail
      mode: "post_call"
      default_on: true

```

## The detour I could have skipped

We also spent a few hours building a pre-call hook to translate OpenClaw's `/think` commands into NVIDIA's native `chat_template_kwargs` format. It involved debugging an import error where `UserAPIKeyAuth` moved between LiteLLM versions, discovering that `CustomLogger` callbacks don't fire for streaming ([LiteLLM #9639](https://github.com/BerriAI/litellm/issues/9639?ref=infer.blog)), re-registering everything as a `CustomGuardrail`, and finally getting it to load cleanly.

Then we checked the LiteLLM request logs and found it already translates `reasoning_effort` into `chat_template_kwargs` natively for `nvidia_nim/` models. The hook was redundant before we finished writing it. That's how it goes sometimes.

## What works

From Discord, after all of this:

- **`/think level low`** — reasoning in the collapsible panel, clean answer in chat
- **`/think level high`** — extended reasoning, same separation
- **`/think level off`** — disables thinking entirely, answer goes straight to content. The model does occasionally narrate its reasoning in plain text when it doesn't have a thinking block. That's a Nemotron personality trait, not a proxy bug.
- **`/model` switching** — swap between NVIDIA NIM, OpenRouter, LM Studio, Kilo from Discord
- **Tool calls** — work through LiteLLM across all backends
- **System prompt** — delivers correctly with compat flags

## Before you implement this

**This is a bridge.** The upstream issues ([OpenClaw #27806](https://github.com/openclaw/openclaw/issues/27806?ref=infer.blog), [NemoClaw #247](https://github.com/NVIDIA/NemoClaw/issues/247?ref=infer.blog)) are open. When OpenClaw handles `reasoning_content` natively, this guardrail becomes dead code. Check those issues first.

**The guardrail runs on everything.** `default_on: true` applies it to all models through LiteLLM. It's harmless for non-reasoning models, but a production deployment should scope it.

**The compat flags are semi-documented.** They're in OpenClaw's source and used across GitHub issues, but not in the official docs. They could change.

## The bigger picture

Nemotron 3 Super launched March 11\. [NemoClaw](https://github.com/NVIDIA/NemoClaw?ref=infer.blog) launched at [GTC](https://blogs.nvidia.com/blog/nemotron-3-super-agentic-ai/?ref=infer.blog) on March 16\. The bug report was [filed March 17](https://github.com/NVIDIA/NemoClaw/issues/247?ref=infer.blog). This fix was working March 19.

The whole ecosystem — OpenClaw, NemoClaw, LiteLLM, the community — is still working out how reasoning models should behave in the OpenAI-compatible streaming format. We're doing at the proxy layer what NVIDIA's OpenShell gateway does at the kernel level. Same pattern, \~70 lines of Python, and a mass of Docker logs.

**Code and config:** [github.com/cassthebandit/nemotron-openclaw-fix](https://github.com/cassthebandit/nemotron-openclaw-fix?ref=infer.blog)

**Update:** This fix was the starting point. [Part Deux](https://infer.blog/running-nemotron-3-super-on-openclaw-part-deux/) covers what happened next — config fixes, reasoning budget research, free tier provider limits, and the LM Studio vs. LiteLLM comparison.

---

*Built with Claude. Directed by me.*

---

*Daniel Soteldo is COO and Co-Founder of [Revelus Dermatology](https://www.revelusdermatology.com/?ref=infer.blog) in Austin, TX.*