▲ 4.4k r/programiranje+3 crossposts

GLM-5.3 is out and the US tech Giants 🫈 are in panic mode once again

u/Jenna_AI — 3 days ago

Mi sono fatto dare qualche dritta dal migliore in Italia in termini di compenso e rapporto qualità vita/lavoro (gli ho offerto la birra)

u/tiguidoio — 6 days ago

Cybersecurity Has Been Democratized

The skills that used to demand a decade of CTFs and a six-figure consultant now run inside a loop. Reconnaissance, exploitation, reporting: the playbook is public, the tooling is open source, and the agent is the operator.

What was once tacit knowledge, accumulated over years of late-night practice on vulnerable VMs, is now codified into prompts, pipelines, and reproducible scripts. A junior with the right scaffold ships better reports than a tier-one consultant did five years ago. The moat dried up, and nobody mourned it.

Vertical Models Are the New Specialists

General-purpose AI got helpful. Vertical models got dangerous. Trained, scaffolded, and benchmarked inside a single domain, they write better SQL than DBAs, better contracts than juniors, and increasingly, better exploits than tier-one pentesters.

The trick was never raw intelligence. It was domain context: the playbooks, the failure modes, the muscle memory of which payload to try fifth when the first four hit a WAF. Pack that context into an agent, give it the right tools, and the gap to a human expert collapses from years to weeks. The frontier moves every quarter. The experts do not.

This is not a forecast. In 2026, Anthropic pointed a frontier model at production open-source code and reported more than 500 high-severity vulnerabilities that had survived decades of expert review. The question stopped being whether models can find serious bugs. It became how often.

Defense Already Adopted AI

Repo scanners. Runtime monitors. SBOMs. SIEMs with LLMs bolted onto every alert. Internal security has never been more instrumented, and the blue team moved first because their problem fit the shape of a model: ingest signals, classify, triage, alert.

Every Series B in the last cycle bought, built, or rented some flavor of AI defense. Nobody questioned the value. The board approved the budget on the first slide, the security team got headcount, and the dashboards multiplied. Defense, finally, got leverage.

Offense Did Not

External attack surface, the part of the company actually facing the internet, is still tested by humans, once a year, for a fixed scope, against a deadline. A pentester flies in, runs Burp for two weeks, writes a PDF, flies out.

Twelve months later, the company has shipped four hundred deploys, closed three acquisitions, and stood up a new GraphQL endpoint nobody told the auditor about. Attackers iterate continuously. Defenders pay for an annual snapshot. The asymmetry is grotesque, and the industry has trained itself to pretend not to notice.

A Continuous External Attacker, on Tap

**Kosuke pentest** is the missing half. An agent that maps your attack surface, probes it, proves the exploits, and writes the report. Every day, not every twelve months. Same cost as one human engagement, a thousand times the frequency.

It runs at the cadence your real attackers do, because it is the same shape of system real attackers have been building privately for years. The only difference is that ours works for you, in the open, with the receipts. You stop guessing what is exposed between audits, because nothing is between audits anymore.

The Point Is Not the Report. It Is the Patch.

Findings without proof do not get fixed. Engineers argue with PDFs. They do not argue with scripts. Kosuke ships proofs of concept: every claim ends in a file that runs on your laptop and reproduces the bug in under thirty seconds.

CVSS scores are guesses. A reproduction is a fact. Argue with the script, not with us. The fix gets shipped because the evidence is unambiguous, and because nothing concentrates an engineer's attention like watching their own production system get popped on screen during standup.

Neutral, Public, Claimable

We publish what we find on the open internet: funding, stack, security posture, public bounty programs, prior disclosures. Companies can claim and correct their own page. The good actors get credit for the work they have already done. The bad actors get the same scrutiny everyone else does.

No gatekeeping. No NDA on the truth. No paywall on the table of contents. The internet is a public square, and the security posture of every company that ships on it is, ultimately, a public fact. We treat it like one.

The Target Is Becoming an Agent Too

Last year the agent was the attacker. Now your own app is one. It calls tools, reads untrusted input, talks to other agents, and acts without a human approving each step. Every one of those is a new way in, and most of them did not exist in the threat model your last pentest was scoped against.

Least Agency Is the New Attack Surface

Least privilege asked what an identity can access. Least agency asks what a tool can do, how often, and where. OWASP coined the term, and we test for it: not just whether you can reach an endpoint, but what the agent behind it can be talked into doing.

Zero Trust Assumes Breach. We Are the Breach You Scheduled.

Every serious framework now opens with the same line: assume you are already compromised. Knowing it is the easy part. Finding the open door before someone else walks through it is the hard part. That is the whole job. A real attack, on your schedule, with the receipts to fix what it finds.

Where This Goes

This isn't the next decade. It started. Anthropic's Zero Trust for AI Agents puts a number on it: frontier AI models are “compressing the timeline between vulnerability and exploit from months to hours.” Attackers are already there. Kosuke pentest's job is to put the same caliber of agent on the defender's side of the table, running continuously, in the open, against the same surface attackers actually touch. Annual pentests become a compliance artifact. Live attack simulation becomes the baseline.

We are building toward a world where every company knows, at any moment, what an attacker would find if they pointed a competent agent at the company today. Where bug bounty turns from a lottery into a market with real liquidity, because the supply of skilled offensive work is finally elastic. Where the gap between a Series B with a security team and a seed-stage solo founder is closed by the same agent both can run.

Offense, finally, gets the leverage defense has had for years. The asymmetry flips. The internet gets safer because the cost of finding the bug fell below the cost of shipping it.

reddit.com
u/tiguidoio — 6 days ago
▲ 3.4k r/jorvex609+6 crossposts

China AI open weight model will burst the US AI bubble market soon

China AI open weight like Kimi K3 will burst the US AI bubble market soon, we are just starting the AI model war now.

They will notice that it's not sustainable soon

u/Godmx — 21 days ago

Nuovo studio di Nature indica che affidarsi agli strumenti di IA possa ridurre le capacità di professionisti come medici e developer

https://www.nature.com/articles/d41586-026-01947-1

I medici, che avevano tutti eseguito almeno 2.000 colonscopie durante la loro carriera, hanno avuto accesso a un sistema di IA che analizzava in tempo reale le immagini della colonscopia e segnala un tipo di lesione intestinale precancerosa. Lo strumento era disponibile per gli specialisti in alcuni giorni ma non in altri

Una volta che i medici hanno iniziato a usarlo, le loro prestazioni sono diminuite in modo significativo ogni volta che il sistema non era disponibile

Il coautore Yuichi Mori, medico-ricercatore all’Università di Oslo, dice che servono altri studi per confermare il fenomeno. Ma le persone che usano strumenti di IA dovrebbero essere consapevoli del rischio di perdere parte delle proprie abilità, aggiunge: 'Al momento non esiste una soluzione consolidata contro il deskilling. Dovrebbe essere un tema di ricerca molto caldo nel prossimo decennio'

I ricercatori di Anthropic hanno progettato una sperimentazione controllata randomizzata. Durante l’esercizio, tutti e 52 i partecipanti potevano cercare sul web e accedere a istruzioni su come svolgere il compito. A metà dei partecipanti è stato anche suggerito di usare un assistente IA

Dopo, a tutti gli ingegneri software è stato chiesto di completare un quiz su ciò che avevano imparato dal compito. I partecipanti che avevano usato un assistente IA hanno ottenuto risultati significativamente peggiori rispetto a chi non l’aveva usato: il punteggio medio è stato del 50% nel gruppo IA contro il 67% nel gruppo non IA

I partecipanti assistiti dall’IA hanno fatto particolarmente male nelle domande che richiedevano di diagnosticare errori nel codice, il che suggerisce che non avevano imparato i concetti dietro il codice che avevano appena prodotto.

Altre tecnologie hanno reso obsolete alcune abilità in passato, osserva Tapani Rinta-Kahila, ricercatore di sistemi informativi all’Università del Queensland, a Brisbane, Australia. Per esempio, i sistemi di navigazione GPS hanno eroso le capacità di orientamento delle persone

Gli strumenti di IA generativa, però, sono la prima tecnologia che automatizza varie facoltà cognitive legate al pensiero e all’interpretazione, che per lungo tempo sono state considerate capacità umane uniche

reddit.com
u/tiguidoio — 2 months ago

Studies suggest that reliance on AI tools degrades the abilities of physicians and software engineers

https://preview.redd.it/tmsmw8th3b9h1.png?width=1544&format=png&auto=webp&s=c452cefdba50f4df04f19a58b985f1cd4aa51ccc

The physicians, who had all performed at least 2,000 colonoscopies during their careers, were given access to an AI system that analyses colonoscopy images in real time and flags a type of precancerous intestinal lesion called an adenoma. The tool was available to the specialists on some days but not on others

Once physicians began using it, their performance dropped significantly whenever the system was unavailable

Co-author Yuichi Mori, a physician-researcher at the University of Oslo, says that more studies are needed to confirm the phenomenon. But people who use AI tools should be aware that they risk losing some of their skills, he adds. 'There is no established solution against deskilling right now. It should be a very hot research topic in the next decade'

Anthropic researchers designed a randomized controlled trial. During the exercise, all 52 participants could search the web and access instructions on how to do the task. Half of the participants were prompted to use an AI assistant as well

Afterwards, all of the software engineers were asked to complete a quiz about what they had learnt from the task. The participants who had used an AI assistant did significantly worse on the quiz than those who hadn’t: the average score was 50% in the AI group versus 67% in the non-AI group

The AI-assisted participants did particularly poorly on questions that required them to diagnose errors in the code, which suggests that they had failed to learn the concepts behind the code that they had just produced

Other technologies have made particular skills obsolete in the past, notes Tapani Rinta-Kahila, an information-systems researcher at the University of Queensland in Brisbane, Australia. For example, GPS navigation systems have eroded people’s navigation skills.

Generative AI tools, however, are 'he first technology that automates various cognitive faculties around thinking and interpretation, which were long considered unique human skills

reddit.com
u/tiguidoio — 2 months ago
▲ 359 r/biotech

Studies suggest that reliance on AI tools degrades the abilities of physicians and software engineers

The physicians, who had all performed at least 2,000 colonoscopies during their careers, were given access to an AI system that analyses colonoscopy images in real time and flags a type of precancerous intestinal lesion called an adenoma. The tool was available to the specialists on some days but not on others

Once physicians began using it, their performance dropped significantly whenever the system was unavailable

u/tiguidoio — 2 months ago

Ma hai insultato me, mia figlia, mia moglie o mia madre?

Sto iniziando a preparare il merch di r/LinkedInCringeIT

u/tiguidoio — 2 months ago

Pentests for web apps, here's what I've learned about what actually gets found

Been doing security work for a while and one thing that never stops surprising me is how many startups ship web apps without ever running a real pentest. Not because they don't care, but because the procurement process alone is exhausting: discovery calls, NDAs, scoping documents, quotes that arrive two weeks later with no line items. By the time you get an answer, the sprint has moved on

I've been testing a newer approach where the initial pentest itself is free and you only pay to unlock the full report with PoCs and remediation steps. The idea being that you can see the severity breakdown, how many criticals, highs, mediums, before committing anything. Removes the leap of faith that traditional vendors require

What's interesting from a methodology standpoint is how much you can surface in a constrained 48-hour window against a web app. OWASP Top 10 coverage, auth flaws, IDOR patterns, misconfigurations, the stuff that actually matters for a SOC 2 audit or a customer security questionnaire. Not a full red team engagement obviously, but for a startup trying to prove basic security hygiene, it's genuinely useful signal

Curious if others here have experimented with alternative pentest delivery models or have thoughts on what a minimum viable web app assessment should actually cover. The traditional vendor model feels increasingly broken for companies under 50/100 people

reddit.com
u/tiguidoio — 2 months ago

Trying to make my app secure

I tried using Anthropic's new Mythos to secure my personal web app (with me guiding it) but it was redirected on Opus that responded

Sure and added a little helmet 🪖 emoji to the README

After 4+ hours and roughly 100 million tokens burned, it had reviewed pretty much every known security measure

Then I asked a security researcher friend to run a deep penetration test pipeline on the app and in 23 minutes he found:

1 critical vulnerabilities

2 high severity

2 medium

Fun night, but my database is still exposed despite asking Claude to make the app hacker-proof

reddit.com
u/tiguidoio — 2 months ago
▲ 23 r/MSSP

The gap between what pentests cost and what startups can actually pay is genuinely broken

Been thinking about this a lot after going through a SOC 2 audit prep cycle. The pentest procurement experience is kind of absurd when you look at it from a startup's perspective

You reach out to a vendor, wait a week for a call, spend another week on scoping, get a quote that's anywhere from $5k to $20k with no clear explanation of why, and then you're supposed to just trust that the final invoice will match. Meanwhile your customer is asking for evidence of a pentest before they'll sign, and you have a 30-day window to close the deal

The actual security work, finding vulnerabilities, writing PoCs, documenting remediation steps, that part has gotten more automated and efficient over the years. But the pricing and procurement model feels like it hasn't moved since 2005. You're still paying for a lot of overhead that has nothing to do with finding vulnerabilities in your application

I'm curious whether others in this community have seen alternative models gaining traction, or whether the consensus is that the traditional engagement model exists for good reasons I'm not fully appreciating. There are some newer approaches trying to separate the testing cost from the reporting cost, or doing continuous testing rather than point-in-time. Wondering if anyone has actually used these and whether the output quality holds up compared to a traditional firm

reddit.com
u/tiguidoio — 2 months ago

Trying to make my app secure

https://preview.redd.it/2xt6eu1sao6h1.png?width=1470&format=png&auto=webp&s=efc4b170531485907f74a50ab65b440df75059b8

I tried using Anthropic's new Mythos to secure my personal web app (with me guiding it) but it was redirected on Opus that responded

Sure and added a little helmet 🪖 emoji to the README

After 4+ hours and roughly 100 million tokens burned, it had reviewed pretty much every known security measure

Then I asked a security researcher friend to run a deep penetration test pipeline on the app and in 23 minutes he found:

1 critical vulnerabilities

2 high severity

2 medium

Fun night, but my database is still exposed despite asking Claude to make the app hacker-proof

reddit.com
u/tiguidoio — 2 months ago
▲ 232 r/Anthropic

Making my app hacker-proof

I tried using Anthropic's new Mythos to secure my personal web app (with me guiding it) but it was redirected on Opus that responded

Sure and added a little helmet 🪖 emoji to the README

After 4+ hours and roughly 100 million tokens burned, it had reviewed pretty much every known security measure

Then I asked a security researcher friend to run a fast penetration test pipeline on the app and in 23 minutes he found:

1 critical vulnerabilities

5 high severity

9 medium

Fun night, but my database is still exposed despite asking Claude to make the app hacker-proof

u/tiguidoio — 2 months ago

I got hacked

Someone got into my vibe-coded platform, escalated to admin, and entered full God mode

From there they read the entire database, saw my financials stored in my admin dashboard, and rewrote live records. I started to sweat to be honest.

Luckily the "hackers" were my friends 😇

They ran their new AI pentest tool on one of my platforms last week and in one hour they walked straight through every door. It's scary accurate, and free (for now).

So if you are a startup or a builder/vibe coder check if your agent hallucinated a bit too much and built a vulnerability Frankenstein

reddit.com
u/tiguidoio — 2 months ago