LiquidAgent-1.2B
A supervised fine-tune of LiquidAI/LFM2.5-1.2B-Thinking on agent traces from TeichAI/Ox-Alpha-Pi-Traces, with tool results properly included in the training data.
What this is
The base LFM2.5-1.2B-Thinking model is a 1.17B-param hybrid architecture (10 LIV-conv + 6 GQA blocks, 32K context, 65536 vocab) with built-in thinking. This fine-tune adapts it for agentic tool-use: the model learns to emit tool calls, see the tool results, and reason about them in subsequent turns.
Key fix over prior attempts: The original parser for the TeichAI data silently dropped all toolResult messages (the type field is always "message" in the raw data, and tool results arrive as message.role == "toolResult" - a branch that did not exist). This fine-tune uses a fixed parser that recovers 100% of tool turns (verified: 562/562 on a 50-file sample).
Training details
- Base: LiquidAI/LFM2.5-1.2B-Thinking (bfloat16)
- Data: 379 examples, ~1.34M assistant tokens (from Ox-Alpha-Pi-Traces, 400 session files, 397 usable)
- Optimizer: AdamW 8-bit, lr 2e-6 to 8e-6 (cosine warmup over 47 steps)
- Hardware: NVIDIA RTX 5090 (32 GB), gradient checkpointing, batch size 1
- Duration: 81 seconds (47 steps)
- Final loss: 0.5965 (step 40) - training completed at step 47
Architecture
| Parameter | Value |
|---|---|
| Params | 1,170,340,608 (tied embedding) |
| Layers | 16 (10 conv + 6 full attention) |
| Hidden size | 2048 |
| FF dim | 12288 |
| Attention heads | 32 Q / 8 KV |
| Vocab | 65,536 |
| Context | 32,768 (max_position_embeddings: 128,000) |
| Dtype | bfloat16 |
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained(
"Compactbot/liquidagent-1.2b",
torch_dtype="bfloat16",
device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained("Compactbot/liquidagent-1.2b")
messages = [{"role": "user", "content": "Write a Python function to parse a JSON config file."}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=512, temperature=0.7, top_p=0.9)
print(tokenizer.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
Honest limitations
- Trained on only 379 examples (~1.3M tokens) - this is a very light adaptation, not a full agent training run.
- The model retains the base's thinking pattern (emits thinking blocks before responding).
- No independent benchmark evaluation was run; quality was verified by generation samples (coherent, on-topic, no degenerate repetition).
- The 32k sequence length requested by the requester was not used in this run (data was packed at shorter lengths to fit the 32GB GPU with 8-bit AdamW).
- Downloads last month
- 515