AI Tinkerers - Post-Training

Filter
The image is a screenshot of a social media post, likely from X (formerly Twitter), featuring a user's profile, a long text post discussing 'evals' in the context of AI and coding companies, and an embedded preview of a video or web page with a transcript.. Text: swyx
@swyx
Subscribe
Claude Code: no evals
[well known code agent company]: no evals
[well known code agent company 2]: kinda halfassed evals
[leading vibe coding company]: no evals
[ceo of company selling you evals]: mmmmm yess all my top customers do evals, you should do evals
[vc's in love with ceo of evals company]: mmmmm yess all my top founders do evals, must do evals
(NOTE: i -do- also think that evals are impt, but the eval pilled ai engineers have also noticed that it is not a strict requirement for success and, at least for 0-to-1 stage, may even be anticorrelated, think thru why)
(4) The future of agentic co...
New Chrome available
Transcript of The future of a...
youtubetotranscript.com
Premium
Search
+ Create
1:00
evals
- Yeah, so the best eval in some sense is
the one that most looks like real life.
And in that case, just using it gives you the
best result.
- We tried really hard, when building Claude
Code, to build a product evals.
- Yeah

Your Conversation Is Out Of Distribution

By Kwindla Hultman Kramer · September 11, 2025 · 8 minutes · 4 photos

Voice agents are hard because most use cases are long, multi-turn conversations and people don’t talk like they write, so LLM performance degrades over turns. Use context engineering + state machines (system prompt, conversation summary, limited tool list, next states). For confidence: do earnest manual testing, then lightweight/repeatable evals. I benchmarked a 30-turn voice scenario with ~75k-token context: every model had errors.

A screenshot of a dark-themed integrated development environment (IDE) displaying Rust programming code, with line numbers and syntax highlighting. On the right, a communication thread is visible, discussing the code's indentation levels and an embedded image of code.. Text: alex
1 month ago in # random - image.png
executor.rs
task_attempts.rs
claude.rs
amp.rs M
db.sqlite
backend > ~/bloop/vibe-kanban/backend/src/executor.rs tor for AmpExecutor > normalize_logs > [@] proce
26 impl Executor for AmpExecutor {
91 fn normalize_logs(
130 let processed = if let Some(msg_type) = json.get("type").and_then(|t|
entries.push(
timestamp: None,
entry_type:
NormalizedE
tool_nam
.to_
action_t
}
Thread
alex Jul 10th at 3:34 PM
amp.rs has 20 levels of indentation
image.png
1 reply
Louis Knight-Webb Jul 10th at 3:49 PM
Probably the most AI generated part of the codebase
Reply...
Also send to # random

70K Lines Later: What We Learned Vibe-Coding a Real App

By Louis Knight-Webb · August 19, 2025 · 4 minutes · 4 photos

Vibe Kanban showed how to ship a real CLI-built app: we capped PRs to <400 LOC and <10 files with a dedicated planning step (affected files, schema changes, test plan), used ts-rs to enforce shared Rust↔TS types via CI regeneration, ran agents in fresh per-session SQLite dev environments, and enforced separation of concerns with Rust crate boundaries (API can’t import SQL/git).

The image displays a conceptual diagram or flowchart illustrating the process of a 'Coding Agent'. It shows a sequence of steps including getting a task, adding it to a list, performing the task, reflecting on the output, and finally producing an output, with some feedback loops.. Text: Coding Agent
Get task
Add to task list
Do Task
Reflect on output
Output

Claude Code: Tips and Tricks

By Anand Tyagi · August 14, 2025 · 22 minutes · 9 photos

Claude Code is an agentic CLI that loops through task → subtasks → execution → reflection, producing files, text, and bash commands; to get strong results, use Anthropic’s Explore/Plan/Execute/Commit, keep context clean with `CLAUDE.md` (root + parents, plus subfolders), and reuse sessions via `/resume` and double escape (Esc+Esc); power up with custom slash commands, subagents, and hooks, plus MCP servers and GitHub integration via `install-github-app` and `@claude`.

A digital illustration depicting a four-step process flowchart titled 'Creating an AI Agent Dynamically'. It outlines the stages from a user request to the deployment of a custom AI agent.. Text: Creating an AI Agent Dynamically
1
User Request
User submits a request for a specific AI agent.
2
Request Analysis
AI analyzes the request and selects necessary tools.
3
Authentication
User authenticates tools for secure access.
4
Agent Deployment
Custom AI agent is generated and deployed for use.

We Built the Lovable for AI Agents - Here’s How

By Karan Vaidya · August 07, 2025 · 11 minutes · 3 photos

We built Lovable for AI Agents because agent building gets stuck on framework choices, tool/auth integration, model selection, and orchestration—so we skip that and generate a working agent in under 10 minutes. Using Composio’s managed tools (execution + authentication for 500+ apps), the app analyzes your use case, selects tools (e.g., Gmail Fetch Emails + Notion Create Page), creates connections via `/api/create-connection`, then dynamically generates a custom frontend + Next.js backend via `/api/generate-agent` for instant deployment.

Post-Training Fellows

Apply to become a Fellow

Have a bold AI idea, credible proof, and a hard-won lesson serious builders can use? Pitch a practical Deep Dive for 130K+ leading AI practitioners globally.

Apply to become a Fellow