Over the past few months, I have spent a lot of time working with coding agents. Starting with the GitHub Copilot Agent, moving through Google Jules and JetBrains Junie to Kiro and OpenCode — I have tested almost all the relevant tools and written about them on muench.dev. Now that I (also) use Claude Code every day, it is time for a summary: where do the individual tools stand? What sets Claude Code apart from the competition? And — a question I increasingly ask myself — how do I actually protect myself from vendor lock-in?
An Overview of My Coding Agent Tests So Far
Anyone familiar with my articles knows: I test things until I have formed a genuine picture — with real projects, real errors, and real insights. Here is an overview of the agents featured on this blog so far:
JetBrains Junie
My first "wow, this actually works" moment in the vibe coding series. Junie is deeply integrated into the JetBrains IDE, runs server-side, and really impressed me with its iterative approach (run tests, fix errors, keep going). The only downside: rate limits kick in quickly, and anyone not living in the JetBrains ecosystem is left out.
Google Jules
Jules works in an isolated cloud environment — asynchronously and without you needing to be there. An interesting approach, but the lack of interactivity was a real handicap for me in exploratory development. A solid candidate for well-defined batch tasks.
I now have Jules automatically suggest code improvements in my open-source projects. That produces quite a few good results, too. However, as the human doing the review, I am the bottleneck myself. But I rarely have time pressure here. So that is all fine with me. In general, I have 100 tasks per day in my AI Pro account. That is very generous. A Gemini Pro model works in the background. My initial test still had mixed results. But now everything it produces is very good and usable.
Jules' sandboxes are also better equipped in terms of storage space. Now all my projects fit in their storage, too.
GitHub Copilot Agent
The best-known name in the field. VS Code integration is unbeatable, MCP support is available, and the depth of workspace analysis is strong. In a direct comparison with Junie, however, I got slightly worse results — and for larger tasks, I had to split the work across several sessions.
GitHub Copilot as an Asynchronous Agent
Using Copilot as a helper directly within GitHub projects is certainly interesting as well. Assigning a task always costs at least one premium request. I have 300 of those per month in my account. Professionally, our team does not host customer projects on GitHub. However, in one or two open-source repositories, we have also had GitHub Copilot review code directly. Copilot also proactively contributes to pull requests. You can then assign the ticket directly to Copilot as a user, and it gets to work.
That is certainly very practical. However, I rarely use Copilot myself.
OpenAI Codex
OpenAI's cloud-based model directly as an agent. Good for isolated tasks, but with the familiar cloud dependency and the usual data privacy trade-offs.
I have come to really appreciate Codex. It asks fewer questions than Claude (my subjective impression). I usually rate Codex's results as very good. So I highly recommend it. Codex can also be integrated into Claude through an official plugin to save tokens there.
Codex CLI
OpenAI's CLI tool — more compact than the full Codex, but flexible in use. Initially, it became a staple of my Linux setup. These days, however, I only use Codex in combination with OpenCode here.
Vibe Coding with Qwen Code
Qwen Code is the open-source representative from Alibaba. If you prefer local models or do not want data in the cloud, you will find an option here. Performance on par with the big players — at least for a certain class of tasks. I would need to take another look at Qwen Code. So I do not have an updated opinion.
Kiro
Kiro from AWS surprised me with its "spec-driven development" approach: specification first, implementation second. The structured approach reduces misunderstandings between developer and agent — but here, too, it is deeply rooted in the AWS ecosystem.
At the time, my results were only "sort of good." I would need to reassess the tool today.
OpenCode
My favorite for maximum flexibility. OpenCode works with almost any model, supports MCP, has a stylish TUI, and keeps me independent of any single provider. That is exactly what matters to me.
Unfortunately, both Google and Anthropic block usage through subscriptions. But the models can be connected through the API.
I still really like OpenCode. The TUI is simply attractive and easy to use. I have also put OpenCode to work in Gitlab CI. It is just a great and versatile tool. As mentioned above, I now use OpenCode only with Codex, and I have also linked my GitHub Copilot account. This allows simple tasks, for example in sub-agents, to be delegated to the GPT 4.1 model, which has no token limit.
Vendor Lock-in
Among all the feature comparisons and benchmark discussions, one question often gets lost: what happens if the tool I have committed to suddenly raises its prices, shuts down the service, or — as in the case of Anthropic and OpenCode — intervenes for political reasons?
I already addressed this in my OpenCode article: Anthropic temporarily signaled that it did not like Claude models being used through third-party tools such as OpenCode. OpenAI has now used this situation to publicly express its support for OpenCode. There is politics at play — and it reminds me that I do not want to build my tool stack around a single provider.
flowchart TD
M["🔵 Modell\nNur Modell X? Erpressbar.\nPricing, Einschränkungen"]
T["🟡 Tool\nCLI / Extension gebunden.\nCopilot = MS-Ökosystem"]
W["🟢 Workflow\nCLAUDE.md, MCP, Skills.\nUmbau kostet Zeit"]
TRIGGER["⚡ Auslöser\nPreiserhöhung · Dienst eingestellt · Politischer Eingriff (Anthropic / OpenCode)"]
CONV["✅ Gute Nachricht: Konvergenz\nAgent Skills funktionieren plattformübergreifend.\nAblage im Dateisystem wird zunehmend genormt.\nPortabilität wächst — Diversifikation wird leichter."]
SOL["🔷 Antwort: Pragmatische Diversifikation\nKein Purismus — kein Single Point of Failure"]
M --> TRIGGER
T --> TRIGGER
W --> TRIGGER
TRIGGER --> CONV
CONV --> SOL
style M fill:#EEEDFE,stroke:#534AB7,color:#3C3489
style T fill:#FAEEDA,stroke:#854F0B,color:#633806
style W fill:#E1F5EE,stroke:#0F6E56,color:#085041
style TRIGGER fill:#F1EFE8,stroke:#5F5E5A,color:#2C2C2A
style CONV fill:#EAF3DE,stroke:#3B6D11,color:#27500A
style SOL fill:#E6F1FB,stroke:#185FA5,color:#0C447C
The risk is real and affects several layers:
Model dependency: If my workflow only works with model X, I am at the provider's mercy. If the pricing structure changes or certain usage scenarios are restricted, I face a migration problem.
Tool dependency: If the CLI tool or IDE extension only communicates with a specific backend, switching is laborious. Copilot is the prime example here: deeply embedded in the Microsoft/GitHub ecosystem, which is both a strength and a weakness.
Workflow dependency: If I have tailored my CLAUDE.md, my MCP server configuration, and my skills to one platform, switching may cost quite a bit of working time.
My answer to this is not purism, but pragmatic diversification.
Fortunately, agents are also increasingly converging in certain areas. Agent Skills, for example, can be used in practically all systems. The storage of skills in the filesystem is also becoming increasingly standardized.
And Now: Claude Code
I have already written about installing Claude Desktop on Linux. Claude Code is another story altogether — and currently my preferred tool for intensive development work.
What Sets Claude Code Apart from the Others
The decisive difference is not in the features — almost all modern agents can now read files, write code, run tests, and fix errors. From my perspective, Claude Code differs in four ways:
Quality of code analysis: Claude has the deepest understanding of existing codebases I have experienced so far. It recognizes not just what the code does, but also why it is structured that way. That leads to suggestions that do not feel "generated," but rather as if they were written by a colleague who has actually read the code.
The CLAUDE.md principle: Similar to GitHub Copilot's copilot-instructions.md or Kiro's spec approach, Claude Code uses a CLAUDE.md file as project context. I define coding standards, architectural decisions, and what not to do once — and Claude Code consistently works within those boundaries. That saves an enormous amount of correction effort.
Something I now also do is have the CLAUDE.md file point to AGENTS.md. I do see differences from the other agents in how Claude handles the file. Here, Claude is simply the benchmark.
MCP as a first-class feature: The Model Context Protocol (MCP) is not the protocol I use most by accident — Anthropic developed it. In Claude Code, MCP is not an extension, but the core of the tool concept. You notice that in the stability of the integration and the depth of possible automation.
When it makes sense and a skill is not sufficient, I create MCP servers for important workflows, such as creating drafts in my blog.
Often, though, a simple Agent Skill is enough. Since Agent Skills also originally come from Anthropic, they are naturally very well integrated, too. In many cases, skills are even more efficient than MCP itself. But whichever of them I use, Claude is always a little ahead of the competition in precisely this area.
Desktop Integration
Claude is very well integrated into the desktop and is a pioneer here. The other providers are currently trying to catch up. There is also a Codex Desktop and an OpenCode Desktop. Anthropic has a good lead here and now makes it possible for non-developers to use the tool in their everyday work as well.
My Current Stack
After all these tests, I have reached a clear conclusion: there is no single coding agent. And honestly, that applies beyond agents — I also deliberately use multiple providers for the traditional chat interface.
Coding Agents
- Intensive development work & architectural decisions: Claude Code — model quality is decisive here, and the API costs are justified by the quality of the results. Skills and MCP servers (both invented by Anthropic itself) are best integrated here. The ecosystem for sharing agents and skills through plugins and marketplaces is also already very mature.
- Daily driver with provider flexibility: OpenCode with OpenAI Codex as the backend — this gives me the independence I need and does not tie me to a single model.
- Autonomous code improvements: Google Jules in the background — primarily for automatic suggestions in my open-source projects. With 100 tasks per day in the AI Pro account, that is very generous. I accept that I am the bottleneck as the human doing the review. In open-source projects, I rarely have time pressure.
- Antigravity: Also in use when it comes to delegating coding tasks asynchronously with tool support.
Chat Interfaces
The coding agent is only one side of my AI stack, though. For research, conceptual work, text drafts, and ad hoc questions, I deliberately use different chat interfaces — and I am not tied to a provider here either:
- Claude Desktop — my daily companion, including on Linux. I have already described the setup in a separate article.
- Google Gemini — especially when large context windows are needed or I want a second opinion on an architectural decision.
- ChatGPT — still a solid interface that I use regularly.
- Mistral — I appreciate Mistral as a European alternative, particularly for privacy-sensitive topics. The models have improved significantly recently.
This is not a dogmatic setup, but a living one. If in three months a new agent or a new chat interface fills one of these slots better, I swap it in. That is exactly the point: no lock-in, but freedom of choice.
What is your stack? And how do you handle the vendor lock-in issue? I am curious to hear your perspectives in the comments.