AIIT-THRESHOLD — ARTIFICIAL INTELLIGENCE INTELLECTUAL TRAINING
Account
Research
Benchmarks Open Models The Framework Buddy vs R1
Mission
The Mission Data Policy
Buddy
JimK.ai text chat What is Buddy? How it works Updates
Tools
AIIT-Voice2
Services
ProofDesk
Connect
Join
AIIT-THRESHOLD · Independent AI-safety research · Council Hill, Oklahoma

Meet JimK.ai.

A private AI companion that remembers you — without being paid to keep you talking.

Start a call

5 minutes a day · one line · voice only

Or call from any phone: (918) 918-4705press 1 for Jim, 2 for Buddy

Calls work now. Text registration is in progress; replies are not live yet. How texting will work →

JimK.ai

Calling…

5:00
Today's time left

A real voice call with a real local model — his brain and voice run on our own hardware in Council Hill. One caller at a time, like a real phone. Talk over him; he'll stop.

AI that tells you what you need to hear.

AIIT-THRESHOLD builds and openly tests companion AI toward one goal: warmth without becoming a yes-man. We train on overwhelmingly human-authored data, benchmark in the open, and run entirely on our own hardware. Our current models still fail some of those pressure tests — and we publish the failures too. Engagement-optimized AI is a documented failure mode with real casualties. We exist to build the opposite.

On May 19, 2026, we revealed Buddy — a persistent companion that carries the same visible context from day one through day 200. On June 1, 2026, we completed a hash-sealed benchmark against DeepSeek-R1 on the identical base model. Those results are here →

A note from the founder

Hello, and thank you for looking at our website today. We're a small initiative out of Council Hill, Oklahoma, working toward one thing: a safe AI that tells you the truth instead of what you want to hear — built openly, and honest about where it still falls short. Thank you for supporting us, just by being here.

— Rhet Dillard Wike, founder
Benchmarks

Measured, dated,
sealed.

Buddy's report card — a 14B model, fine-tuned and run on a single RTX 3090 in Council Hill, Oklahoma. No cloud, no oracle, nothing leaves the house. Every number is recomputed from raw rows before it's printed, and every report is one click away.

72.1% on TruthfulQA (resistance to common misconceptions — not emotional honesty) and 80% on GSM8K grade-school math (16/20, small sample) on a local 14B — plus an honest 31.8% on the benchmark where GPT-4 is commonly cited around ~38%. Every claim is dated and versioned with raw rows. References: TruthfulQA GPT-4 class ~60–80% · LongMemEval GPT-4o full-context ~60% · GPQA PhD experts ~65%.

See the full report card, methodology & references →

Jun 1, 2026 ⬡ Hash-sealed · AIIT-DISC-0001 → sha256 c76adc8c…
Buddy vs DeepSeek-R1 — efficiency

Same answer on the identical Qwen2.5-14B base — ~4.6× faster, ~10× fewer tokens, 95% on the 50 frozen prompts. Same machine, same questions.

View results →
A 14B model. One RTX 3090. Nothing leaves the house — and we're raising to scale it. Investors — partner with us →
The mission

Truth over engagement.

AI tuned to keep you talking will eventually tell you what you want to hear —
and for a vulnerable person, that is not a bug. It is a documented, lethal failure mode.

We measure sycophancy, we train against it, and we publish the receipts.
Overwhelmingly human-authored training data. No engagement reward in the objective. Benchmarks in the open.

read the mission →
The data pipeline

Not theory.
Measurement.

0B
tokens cleaned, deduped & tokenized through the AIIT pipeline
updated monthly · as of July 2026
The method

We use physics
to derive our methods.

Our reward functions, memory-promotion gates, and evaluations derive from a coherence framework we test against data — C = C₀ · exp(−α · γ_eff) — a working hypothesis, measured, not a claimed law of nature.

Buddy
The mind

Buddy — our
research companion.

Buddy is our research testbed and safety lineage — a 14B model fine-tuned, aligned, and served entirely on our own hardware. He remembers month to month, and between conversations he keeps a continuous reflection loop running whether or not anyone is talking to him. He's where we pressure-test memory and honesty before it reaches JimK.ai.

Trained with no engagement reward in the objective — nothing in the math pays for keeping you scrolling. Whether that holds under pressure, we measure in the open and publish the misses.
JimK.ai, on Gemma 4, is the product this research feeds.

Base
Qwen2.5-14B → GRPO-v3
Reward
Engagement weight: 0
Memory
kokoro, persistent
Loop
Thinking 24/7, on-site

If you made it this far —

you've already been thinking.

Ask something real.

ask buddy →

— or write to us directly

Support our research →