Benchmarking OpenTelemetry: can AI trace your failed login?
A lot of vendors pitch AI SRE. We tested 14 models across 11 programming languages; even the best ones struggle with instrumenting code with the leading open-source standard, OpenTelemetry.
Insights on agentic coding tools, LLM evaluation, benchmarking, and simulation environments.
Stay tuned for future posts and releases
Subscribe via RSSA lot of vendors pitch AI SRE. We tested 14 models across 11 programming languages; even the best ones struggle with instrumenting code with the leading open-source standard, OpenTelemetry.
Prompts are specs, not code. This influences git workflows for vibe coding, tracking LLM prompts in GitHub repositories, managing commit messages, and debugging non-deterministic AI outputs.
AI reasoning models like DeepSeek-R1, agentic coding tools like Claude Code, and image generation with Nano Banana Pro set daily software engineering standards.

Standardizing AI agent evaluation with Harbor: an open-source framework for reproducible benchmarks, reinforcement learning, and collaborative evals.

Comparing Google Antigravity and Claude Code for AI-assisted workflows, and why custom Claude Skills might be the better approach.

A reality check on AI enthusiasm: how a simple chart generated vitriolic reactions outside the tech bubble.
Finally, an AI that can draw a map or create an infographic. The capability of leveraging tools pushed the frontier of image generation.

We had a great team, $2.5 million, and a validated problem. A year later, we sold our IP for parts. Here’s what we learned about urgency and co-foundership.
Local LLMs prioritize privacy over security. Our research reveals a 95% backdoor injection success rate.
AI excels at clean algorithms but fails at messy, real-world codebases. The solution lies not in Go-like intelligence, but StarCraft-like complexity.
We tested 19 LLMs on their ability to handle real-world software engineering tasks like compiling old code and cross-compiling. See how Anthropic, OpenAI, and Google models stack up in our new benchmark – CompileBench.
We expected small models to be fast, but our benchmarks revealed a common reliability trap. Here’s our deep dive on finding and fixing it.
A one-line command to diagnose server health. Uses Nix to fetch tools without sudo and an LLM to summarize the output. No installation required.
Deep dive into the Tau² benchmark that goes beyond LLM evaluation to reveal innovative methodologies for testing AI agentic systems in realistic scenarios. Learn how this framework can transform how we test AI-powered software.
Learn how Go practitioners ship telemetry in 2025 – what works, what hurts, and the tools, workflows, and guardrails they rely on for metrics, traces, and logs.