> ## Content Index
> Fetch the complete content index at: https://infer.blog/llms.txt
> Use this file to discover other available public pages before exploring further.

# Running Nemotron 3 Super on OpenClaw: Part Deux
- URL: https://infer.blog/running-nemotron-3-super-on-openclaw-part-deux/
- Published: 2026-03-26T15:43:06.000Z
- Updated: 2026-03-27T03:09:53.000Z
- Description: A week after the original fix, more bugs, more configs, and an agent writing wrong notes about itself. 987 requests, 35 million tokens, $0.00.
- Author: Daniel Soteldo
- Tags: AI, Nemotron, OpenClaw, LiteLLM, Nvidia NIM

**TL;DR:** A week after [getting Nemotron 3 Super running in OpenClaw](https://infer.blog/running-nemotron-3-super-on-openclaw-a-free-inference-rabbit-hole/), I realized it was time to roll my sleeves up again since my agent was confidently writing wrong notes about her own behavior in system memory files and spiraling with endless reasoning messages. I ended up working through a couple of bugs, made some config improvements, and learned some things worth writing down. All of this was across 987 requests, 35 million tokens, and $0.00... NVIDIA is actually walking the walk on free builder access, which creates real opportunities and pushes the whole industry forward. If you're interested in running Nemotron as a personal agent via OpenClaw instead of through NemoClaw, this article is for you.

---

NVIDIA launched [NemoClaw](https://github.com/NVIDIA/NemoClaw?ref=infer.blog) at GTC on March 16, right around the time the [original post](https://infer.blog/running-nemotron-3-super-on-openclaw-a-free-inference-rabbit-hole/) went up. NemoClaw is an [OpenShell](https://www.nvidia.com/en-us/ai/nemoclaw/?ref=infer.blog) runtime that wraps OpenClaw in a sandbox with policy-based security, audit logging, and network isolation. It's aimed at enterprises who need to get past their security team. It's [early alpha](https://docs.nvidia.com/nemoclaw/latest/about/overview.html?ref=infer.blog), and there's still an open [blank response bug](https://github.com/NVIDIA/NemoClaw/issues/247?ref=infer.blog) from reasoning models and [local model routing](https://github.com/NVIDIA/NemoClaw/issues/385?ref=infer.blog) is still being developed. The LiteLLM guardrail from the original post is a workaround for your OpenClaw while NVIDIA builds out their tooling and OpenClaw is further developed. NemoClaw and this stack solve different layers of the same problem: NemoClaw handles security and sandboxing, this handles the model quirks that make Nemotron work as a daily driver in OpenClaw.

The original post ended with a working fix. I ran it for a few days, let it sit, and came back when I had time to dig in... It wasn't paradise, but eventually, there was trouble.

## How did everything go off the rails in a week?

After swapping Cass (my OpenClaw agent running on Discord) over to Nemotron and LiteLLM, I started experiencing issues after upgrading to the latest version of OpenClaw and setting the think mode to high. Issues started snowballing.

This was from Cass's own memory files — notes she'd written about herself during a nightly self-review:

> *"Pattern observed 2026-03-19: early messages in session produce truncated/mangled output, but model warms up after 5-6 exchanges and produces full responses. Not ready for main session use; investigate warm-start behavior."*

Naturally, I thought this was a warm-start problem, but after [researching the documentation](https://docs.nvidia.com/nemotron/nightly/usage-cookbook/Nemotron-3-Super/OpenScaffoldingResources/README.html?ref=infer.blog) and digging around in various log files, I landed on a missing config — specifically `max_tokens`, which was too low at 4096\. Turns out, `max_tokens` and `temperature: 1.0` (with `top_p: 0.95`) are all recommended settings in both the [model card](https://build.nvidia.com/nvidia/nemotron-3-super-120b-a12b/modelcard?ref=infer.blog) and [NVIDIA's agentic coding tools guide](https://docs.nvidia.com/nemotron/nightly/usage-cookbook/Nemotron-3-Super/OpenScaffoldingResources/README.html?ref=infer.blog). With only 4096 tokens of headroom, the model was running out of space mid-think before it could close its `</think>` tag. Without that closing tag, the model can't transition from reasoning to responding, so it just sits there. That's what was causing the truncated output and the false "warm-start" pattern Cass had been logging.

**First real fix:** Set `max_tokens: 32768`, `temperature: 1.0`, and `top_p: 0.95` in `litellm_config.yaml`. After redeploying, a quick check of the LiteLLM logs confirmed the new settings were making it through to NIM.

```yaml
model_list:
  - model_name: nemotron-nvidia
    litellm_params:
      model: nvidia_nim/nvidia/nemotron-3-super-120b-a12b
      api_key: os.environ/NVIDIA_API_KEY
      drop_params: true
      max_tokens: 32768
      temperature: 1.0
      top_p: 0.95

```

## thinkingDefault: high — grab a straitjacket

With `max_tokens` fixed to 32768, the model had room to think. But now a different problem showed up.

OpenClaw's `thinkingDefault` is a global setting. Mine was set to `high`. For Anthropic this is fine. Fine if you have some seriously important noodling to do where you need a big brain, or fine if you just don't mind setting tokens and money on fire. For Nemotron through LiteLLM, `high` translates to [chat\_template\_kwargs: {enable\_thinking: true}](https://docs.nvidia.com/nim/large-language-models/latest/reasoning-model.html?ref=infer.blog) with no cap. This means Nemotron will burn its entire 32K output budget on reasoning, never emit `</think>` (since it never decides to stop thinking), and ultimately just produce nothing.

What this looks like in practice: your chat window shows the agent typing, the typing bubble just sits there, and then eventually disappears. If you turn on reasoning visibility, you'll see your agent pontificating over everything and spiraling into oblivion until the 32K window fills up and the interaction ends. It's pretty wild actually — your agent will look nuts talking and arguing with itself, which is amusing for maybe a minute or so.

`medium` sends [{enable\_thinking: true, low\_effort: true}](https://docs.nvidia.com/nim/large-language-models/latest/thinking-budget-control.html?ref=infer.blog) — NVIDIA Inference Microservices' (NIM) bounded mode. The model reasons within a limit, closes out, and responds.

**Second fix:** run `openclaw config set agents.defaults.thinkingDefault medium`.

```bash
openclaw config set agents.defaults.thinkingDefault medium
openclaw gateway restart

```

There's an alternative approach worth knowing about. NIM supports a [reasoning\_budget](https://docs.nvidia.com/nim/large-language-models/latest/thinking-budget-control.html?ref=infer.blog) parameter that hard-caps thinking tokens via a logits processor. The model doesn't always use its full budget since it stops thinking when it reaches a conclusion, so the cap is a ceiling, not a target. [Research shows](https://cobusgreyling.medium.com/nvidia-nemotron-3-super-reasoning-budget-sweep-f600358bbbdf?ref=infer.blog) quality plateaus at 1,024 thinking tokens on math benchmarks, meaning unconstrained thinking can waste up to 16x the compute for zero accuracy gain.

There's a catch though: that testing only covers single-turn math, not agentic tool calling. [Daily.co found](https://www.daily.co/blog/nvidia-nemotron-3-super/?ref=infer.blog) that tool calling performance drops significantly with reasoning disabled entirely because the model uses its thinking trace to verify tool call results. For long tool call chains where each step builds on the last, the right reasoning budget is genuinely an open question. My intuition is it's probably logarithmic, approaching a max, where early tool calls need meaningful headroom but each additional call adds less. But that's speculation.

![Discord showing Cass making 20+ sequential exec tool calls in a single task, searching for SearXNG across env, docker, ps, launchctl, lsof, netstat, curl, find, pip, homebrew, and config files](https://infer.blog/content/images/2026/03/IMG_7118.png)

A single prompt, 20+ tool calls. This is the scenario the research doesn't cover.

I didn't end up using `reasoning_budget` directly because LiteLLM already translates `reasoning_effort` into `chat_template_kwargs` internally, and setting a static `chat_template_kwargs` alongside it might cause a collision. Setting OpenClaw's `/think` to `medium` sidesteps it entirely. The collision risk is documented [in the repo](https://github.com/cassthebandit/nemotron-openclaw-fix?ref=infer.blog) (now updated to v2) if you want to dig into it.

## There is OpenRouter Free and then there is NVIDIA Free

My fallback chain was LM Studio running on a local M4 box → OpenRouter → NIM → Kilo. The M4 doesn't run LM Studio at night, so every overnight cron job and subagent would hit the fallback chain. What I found was that OpenRouter was slower and it was capped at 50 free requests per day, and then OpenClaw would eventually fallback to NIM. This is how I learned that OpenRouter caps their usage to 50 requests. I never thought to look, but it's worth noting since anything more than casual usage will hit Nemotron free tier limits when using OpenRouter pretty quick.

If you're shopping free Nemotron inference providers, here's what I found:

| Provider                                                                                                                                  | Free Tier           | Limit Type  | Notes                                                            |
| ----------------------------------------------------------------------------------------------------------------------------------------- | ------------------- | ----------- | ---------------------------------------------------------------- |
| [NVIDIA NIM](https://build.nvidia.com/nvidia/nemotron-3-super-120b-a12b?ref=infer.blog)                                                   | 40 req/min          | Per-minute  | Developer Program, most generous by far                          |
| [Kilo](https://blog.kilo.ai/p/nvidia-nemotron-3-super-launch?ref=infer.blog)                                                              | Free (limited time) | Promotional | No published cap yet                                             |
| [OpenRouter](https://openrouter.zendesk.com/hc/en-us/articles/39501163636379-OpenRouter-Rate-Limits-What-You-Need-to-Know?ref=infer.blog) | 50 req/day          | Per-day     | Reduced from 200 in April 2025; failed requests count toward cap |

The difference between 40 requests per minute and 50 requests per day is massive.

![LiteLLM usage dashboard showing 987 total requests, 35,136,570 tokens processed, and $0.00 spend across openrouter/nvidia/nemotron-3-super-120b-a12b:free (192 requests) and nvidia_nim/nvidia/nemotron-3-super-120b-a12b (795 requests)](https://infer.blog/content/images/2026/03/IMG_7573.png)

192 requests to OpenRouter, 795 to NIM. The 192 should have been going to NIM the whole time.

**Lesson learned:** Take the time to set up all your model configs properly when it's in front of you; you'll save yourself the time of troubleshooting issues later.

## Two viable paths: LM Studio vs. NIM through LiteLLM

If you're running Nemotron on OpenClaw, there are a couple architectures you can choose between. Each has their own pros and cons.

[LM Studio](https://lmstudio.ai/?ref=infer.blog) runs the model locally and owns the full stack from inference to API response. Their server parses the `<think>` tags directly from the model output, [separates reasoning content from response content](https://lmstudio.ai/changelog?ref=infer.blog) in the OpenAI-compatible API, and reports accurate token counts back to OpenClaw. Everything just works. No guardrail needed, no token counting issues, no reasoning content bleeding into chat.

The NIM through LiteLLM path is different. NIM returns `reasoning_content` as a separate field, but by the time it passes through LiteLLM's translation layer and reaches OpenClaw, there are three pieces of software that need to agree on how that field is formatted. Right now they don't fully agree. That's why the [guardrail](https://github.com/cassthebandit/nemotron-openclaw-fix?ref=infer.blog) from the original post exists, and it's why the token counter breaks on this path.

Nemotron also has a habit of thinking out loud. It'll narrate its reasoning as plain text in the content stream, separate from the `reasoning_content` field entirely. LM Studio's server-level parsing catches this and routes it correctly. Through LiteLLM, it comes through as regular chat messages. You can partially mitigate it by adding instructions to your agent's config files, but it'll still break through from time to time. It's a [model trait](https://lmstudio.ai/llms-full.txt?ref=infer.blog), not a bug in any particular tool.

If you're running Nemotron locally and don't need multi-provider routing, LM Studio is the cleaner path. If you need NIM for the free hosted inference or for the speed, the LiteLLM guardrail path can handle it.

![OpenClaw status bar on LM Studio path showing working token count at 31k/262k (12%), 50% cache hit, cost $0.00](https://infer.blog/content/images/2026/03/IMG_0562_resized.png)

![OpenClaw status bar on NIM through LiteLLM path showing broken token count at 0/262k (0%), 0 compactions](https://infer.blog/content/images/2026/03/IMG_9996.png)

Same model, same OpenClaw. Left is LM Studio, right is NIM through LiteLLM. The token counter tells the story.

## Token counting and context window

There's an [open issue](https://github.com/openclaw/openclaw/issues/50795?ref=infer.blog) with OpenClaw (#50795) where your token count will always stay at 0 when using NIM through LiteLLM. Mine always says `0/262k (0%)`. There's a [related issue (#52181)](https://github.com/openclaw/openclaw/issues/52181?ref=infer.blog) that describes OpenClaw not mapping OpenAI-format usage fields back to its internal format for non-Anthropic model responses — that one was closed, but the behavior still persists on the LiteLLM path.

The practical problem is that OpenClaw uses the token count to decide when to compact your session. Token count reads zero, compaction never fires, sessions grow until the model degrades. During one debugging session, the logs showed context climbing from 152 messages and 301k characters to 176 messages and 321k characters before the agent started stopping mid-task.

I also found a global config cap (`agents.defaults.contextTokens: 200000`) that was overriding every model's actual context window. Backed that out and added individual configs per model. The suggested context window for Nemotron is [262K when using vLLM/NIM](https://build.nvidia.com/nvidia/nemotron-3-super-120b-a12b/modelcard?ref=infer.blog), which is a bump over Anthropic models at 200k. It technically supports up to [1M tokens](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-FP8?ref=infer.blog), but that requires `VLLM_ALLOW_LONG_MAX_MODEL_LEN=1` and risks CUDA OOM on smaller hardware.

For now, I just compact or create new sessions from time to time or if I notice model degradation.

## Session startup was taking two minutes

Every time I started a new session, Cass would sit there for about two minutes before saying anything. Turns out my AGENTS.md file was telling the agent to read SOUL.md, USER.md, MEMORY.md, and AGENTS.md at the start of every session, but OpenClaw already injects all of those into the system prompt before the agent even starts. The instructions were left over from testing MiniMax M2.5, where system prompt injection didn't work properly and the agent had to manually re-read its config files. Different model, different quirk, and configs that I had forgotten about.

On Anthropic each redundant file read takes milliseconds, so it was invisible. On Nemotron with `thinking: medium`, every tool call goes through a reasoning pass at 20-30 seconds each. I updated AGENTS.md so the agent knows those files are already in context and only reads the two daily log files that OpenClaw doesn't auto-inject. Startup went from about six tool calls down to two, roughly two minutes to 30-40 seconds.

## The gap that hasn't been built yet

Working through all of this, I kept running into the same limitation: there's no way to configure different behavior per model in OpenClaw. `thinkingDefault` is global. AGENTS.md is one file. The workspace files are the same regardless of which model is running.

That's fine if you're only running one model. It's a challenge when you're switching between alternative cost-effective models (Nemotron, MiniMax, etc.) and metered frontier models (Sonnet, Opus, etc.). The right startup sequence and thinking config for one is wrong for the other. Right now, you just have to keep the different setups in mind.

There's an [open issue](https://github.com/openclaw/openclaw/issues/6528?ref=infer.blog) asking for per-agent `thinkingDefault`. The 2026.3.23 release added per-agent `reasoningDefault` (controls whether reasoning is displayed) but not `thinkingDefault` (the actual thinking level sent to the model). Still global.

## Where things landed

The model is solid. The stack is stable. Nemotron is a great option on OpenClaw. Be patient with it, and it'll get through most tasks without issue.

987 requests, 35 million tokens, $0.00\. NVIDIA's free builder access is what made all of this practical. It's truly a great time to be a builder.

Code and config: [github.com/cassthebandit/nemotron-openclaw-fix](https://github.com/cassthebandit/nemotron-openclaw-fix?ref=infer.blog)

---

*Claude helped debug the issues and co-wrote this post. I directed the investigation, made the technical decisions, wrote key sections, and did the final edit.*

---

*Daniel Soteldo is COO and Co-Founder of [Revelus Dermatology](https://www.revelusdermatology.com/?ref=infer.blog) in Austin, TX.*

---