Hacker Newsnew | past | comments | ask | show | jobs | submit | sinatra's commentslogin

> it's been playing the game of economic protectionism for centuries.

Equal to or more than other major countries? If so, can you show me anything supporting your claim?


I wasn't making a comparison or a value judgement. We're obviously both aware that other countries (and blocs, like the EU) also play the same game.

Rather, I was simply noting that the US' "freedom" branding that the poster was referring to doesn't necessarily apply in the way they were assuming (their point was that the "free" US restricting models more than "authoritarian" China was ironic) given the US' history of economic protectionism.


I've tried Chinese open models few times before. They were fine, but they didn't come close to the benchmarks they were claiming.

Now, maybe GLM 5.2 is close to Opus 4.7, but I don't wanna keep checking them and keep finding that they're still benchmaxing and aren't at GPT (my choice) or Opus level. The boy who cried wolf, I guess.


Yes, my experience has been the same as yours. I find that the performance of open models is quite acceptable, even good, at one-off questions or small tasks. But they are quite unreliable at long horizon goals.


Hah. And you think a Govt agency will be able to do a rigorous enough examination to eliminate people who don't know that just because a method in C# is asynchronous does not mean it executes out of order?

Reminds me of a friend whose job application got rejected by some Govt agency in Canada due to "experience mismatch." Job required "Software Programmer" experience but he was a "Software Engineer" instead.


Obsidian is great. And glad to see how you’re looking at plugin security. But one more thing you should consider is: How do you reduce the need for plugins for basic product behavior. E.g., I use a plugin to be able to open a file in new tab instead of replacing current tab. That should be a setting, not a plugin I’m forced to use.



Easy to say that, isn’t it? I can almost guarantee that YOU also want and do consume much more than an average human on this planet. Never mind the fact that humans consume more than their share of the planets resources anyway. Whatever the definition of “their share” is.


This is why I used the word troubling, instead of wrong or evil or anything like that.

Yes, the devil is in the details. I just don't think it is unreasonable to be troubled by people wanting to be rich.


Will 1.5B people have a lot of very intelligent people too? Yes, some of the most intelligent! Will those intelligent people have the educational opportunities and research opportunities to be able to use that intelligence to deliver a SOTA model any time soon? Especially with so many resource limitations they face, I doubt it.


Education is no longer locked behind academia. Even elite universities were never really about teaching in the first place and more about connecting rich people. Today everyone with internet access can easily get all the education they need to work in this field.


Oh. So when you say “May we please have terraform back?” You mean “May we please have terraform back at my employer?” Why are you posting such an employer specific request on a public forum?


Because it was meant as a rhetorical device, not a literal request.


Piggybacking on this post. Codex is not only finding much higher quality issues, it’s also writing code that usually doesn’t leave quality issues behind. Claude is much faster but it definitely leaves serious quality issues behind.

So much so that now I rely completely on Codex for code reviews and actual coding. I will pick higher quality over speed every day. Please don’t change it, OpenAI team!


Every plan Opus creates in Planning mode gets run through ChatGPT 5.2. It catches at least 3 or 4 serious issues that Claude didn’t think of. It typically takes 2 or 3 back and fourths for Claude to ultimately get it right.

I’m in Claude Code so often (x20 Max) and I’m so comfortable with my environment setup with hooks (for guardrails and context) that I haven’t given Codex a serious shot yet.


The same thing can be said about Opus running through Opus.

It's often not that a different model is better (well, it still has to be a good model). It's that the different chat has a different objective - and will identify different things.


My (admittedly one person's anecdotal) experience has been that when I ask Codex and Claude to make a plan/fix and then ask them both to review it, they both agree that Codex's version is better quality. This is on a 140K LOC codebase with an unreasonable amount of time spent on rules (lint, format, commit, etc), on specifying coding patterns, on documenting per workspace README.md, etc.


That's a fair point and yet I deeply believe Codex is better here. After finishing a big task, I used two fresh instances of Claude and Codex to review it. Codex finds more issues in ~9 out of 10 cases.

While I prefer the way Claude speaks and writes code, there is no doubt that whatever Codex does is more thorough.


Every time Claude Code finishes a task, I plan a full review of its own task with a very detailed plan and it catches itself many things it didn’t see before. It works well and it’s part of the process of refinement. We all know it’s almost never 100% hit of the first try on big chunks of code generated.


How exactly do you plan/initiate a review from the terminal? open up a new shell/instance of claude and initiate the review with fresh context?


It depends on the task but I have different Claude commands that have this role, usually I launch them from the same session. The command has the goal of doing an analysis and generating a md file that I can execute with a specific command and the md as parameter. It works quite well. The generated file is a thorough analysis of hundred of lines with specific coded content. It’s more precise that my few line prompt and help Claude stay on rails


Yeah. It dumps context into various .md files, like TODO.md.


Thanks for the tip. I was dubious, I tried GPT 5.2 for a start on a large plan and it was way better than reviewing it with Claude itself or Gemini. I then used it to help me with feature I was reviewing, it caught real discrepancies between the plan and the actual integration!


This makes me think: are there any "pair-programming" vibecoding tools that would use two different models and have them check each other?


Have you tried telling Claude not to leave serious quality issues behind?


Let’s call it JoyScript so it still shortens to JS. And so at least the name as some joy in it even if the language doesn’t.


Hah. It can’t be “I need to spend more time to figure out how to use these tools better.” It is always “I’m just smarter than other people and have a higher standard.”


Show us your repos.


My stack is React/Express/Drizzle/Postgres/Node/Tailwind. It's built on Hetzner/AWS, which I terraformed with AI.

It's a private repo, and I won't make it open source just to prove it was written with AI, but I'd be happy to share the prompts. You can also visit the site, if you'd like: https://chipscompo.com/



Spot on.


The tools produce mediocre, usually working in the most technical sense of the word, and most developers are pretty shit at writing code that doesn't suck (myself included).

I think it's safe to say that people singularly focused on the business value of software are going to produce acceptable slop with AI.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: