How much can I trust a cardboard mock-up for an under-sink RO?

I'm trying to judge the space around my disposal before ordering anything. My plan is to make a cardboard box the size of the unit and set it in the sink cabinet. That should show whether the door hinge hits it and how much working room remains.

The Viomi V8 Pro is listed at 16.3 inches deep, 5 inches wide, and 12.8 inches high, but its manual suggests reserving a space of 20 inches by 12 inches by 20 inches under the sink. I would also tape that larger outline inside the cabinet instead of treating the box as the whole fit check. If you've tried this, what did the mock-up still miss once the tubing and fittings were in place?

reddit.com
u/Any-Farm-1033 — 7 days ago
▲ 1 r/PFAS

How do you verify a retail RO model's PFAS claim?

I bought a Viomi Vortex 8 Pro, and after it arrived I checked the model number against Viomi's V8 Pro certificate page and the performance information for the filter. Seeing the model-specific report reference and the PFAS/PFOA/PFOS testing language line up was reassuring. I still wanted to understand which standard and contaminant each statement covered instead of treating every standards reference as interchangeable.

The way I am checking it is fairly unglamorous: copy the full retail model number, find the performance data sheet, look for the exact contaminant and standard, then see whether a public certifier listing uses the same identifier. A product title or a report for a different code is not enough.

The EPA home-filter guide points people to the specific PFAS claim under NSF/ANSI 53 or 58 and to the certifier's directory. It also warns that an existing certification does not automatically answer what happens under newer EPA limits.

EPA guide:
https://www.epa.gov/cleanups/reducing-pfas-your-drinking-water-home-filter

Viomi V8 Pro certificate page:
https://water.viomi.com/pages/certificate-v6-8-pro

NSF drinking-water treatment unit directory:
https://info.nsf.org/Certified/dwtu/

If you review these documents often, what identifier do you check first to confirm that the retail name and certification listing refer to the same unit?

u/Any-Farm-1033 — 13 days ago
▲ 0 r/stocks

War headlines added 82 points. One export print erased 53, and the market took it all back by Thursday. That round trip says it all.

I spent the weekend after July 8 thinking the energy cost channel would dominate this tape for more than two sessions. Brent had jumped about 5% to around $76 when the waiver was revoked on July 7, then settled another 9.5% higher on July 13 alone as Iran's retaliation against US bases in Kuwait and Bahrain widened the escalation. Shanghai closed Friday July 10 at 3,996.16, down 1.00%, and I assumed we were in for a drawn out repricing of manufacturing margins.

Monday July 13 proved me partly right in direction, wrong in duration. The index closed at 3,913.79, down 2.06%, a three month low. The selling was broad, more than four thousand names down on the session, and the export facing industrial complex was not spared. Higher energy input costs, tighter industrial margins, the standard playbook.

Then Tuesday July 14 happened. China's June trade data printed: exports up 27% to a record $412 billion, led by AI hardware. Shanghai closed at 3,967.13, up 1.36%. That single session regained roughly two thirds of the 82 point gap between the July 10 and July 13 closes. Oil did not collapse to allow this. It pushed to a one month high that same day. The market simply decided the demand side evidence outweighed the energy cost fear, and it decided it inside a single session.

This is where I need to be honest about what I cannot separate cleanly. The trade print landed on a tape that had just made a three month low. Some of that 53 point recovery was the data mattering. Some of it was selling exhaustion. I cannot cleanly apportion the two, and anyone who claims they can split it cleanly is guessing.

I also need to flag the living risk. Oil remains elevated after July 13. The manufacturing cost squeeze that the selling priced has not resolved. If those input costs persist into third quarter margins, the July 14 bounce will look premature. The structural claim I am making, that this market prices demand side evidence over energy cost fear, fails if the next oil headline outweighs the next data print.

But the speed asymmetry was the tell that kept me up at night. Two sessions to price a war driven cost shock. One session to undo most of it on a trade number. That ratio suggested the underlying bid was organized around end demand visibility, not input cost anxiety. The energy sensitive industrial names are the ones paying the higher input bill. The AI hardware exporters filling record orders are the ones the trade print spoke for.

Then Wednesday July 15 tested it. Shanghai closed near 3,955, down 0.30%, the day the US restored its naval blockade on vessels traveling to and from Iranian ports. Thursday July 16 finished the round trip. Shanghai closed at 3,882.41, down 1.85%, a fresh three month low below the July 13 close, in a broad tech selloff that also took the Kospi down more than 6%, while Brent eased to about $84.6 after four straight up sessions. The entire 53 point July 14 recovery was given back and more.

The one session demand side repricing on July 14 was real, but it did not survive the escalation from a price shock to a supply blockade. The July 16 leg was as much a global chip and valuation rout as an oil trade, so the clean 'demand beats cost fear' read failed its first live test.

The ratio said demand wins. The next two sessions said that only holds until the supply risk stops being a price and becomes a blockade. I am keeping the ratio, but I no longer trust it unconditionally. No direct China index position here, this is observation.

I keep coming back to the fact that I am long both sides of this argument in practice, not by design but by the shape of the index. A China manufacturing book that contains both the energy cost losers and the AI demand winners of exactly this tug of war partly nets the trade inside one wrapper: CNQQ carries CATL at about 6.4% and Zhongji Innolight at about 6.4% per its May 2026 published holdings, so the battery name paying the higher energy bill and the optical module name filling the record export orders sit at equal weight, which you can read as diversification or as muddle depending on which leg of this week's round trip you believe. KWEB is internet only and holds zero A shares, so it captures neither side of this particular trade. I should note the obvious caveats: the fund is tiny, about sixteen million in AUM last I checked, and launched only in September 2025, so the live history is thin.

reddit.com
u/Any-Farm-1033 — 29 days ago

Most "best AI tool" ranking pages are SEO funnels, here is how to spot them

There are too many "best AI tool" ranking pages now. After getting fooled by a few, I started paying attention to the business model behind these pages rather than just the content.

The page exists to rank, not to test. A real benchmark starts with a methodology and produces rankings as output. A funnel starts with the desired ranking and builds a page around it. You can usually tell by checking whether a methodology section even exists, and if it does, whether it actually constrains the results.

The reviewer is the product. Some ranking sites are run by content creators who also do sponsored work for the tools they rank. That does not automatically mean the results are wrong, but when a reviewer regularly does sponsored content or paid collaborations with their top-ranked tool and does not disclose this on the ranking page itself, you should be skeptical.

One tool dominates every category. Real-world tools have tradeoffs. Speed versus quality, price versus features. When one tool sweeps across the board, the ranking is telling you more about the ranker than the tools.

Affiliate links with no disclosure. Monetization is fine, hiding it is the tell. Honest pages disclose their financial relationships. Funnels bury them.

"Updated 2026" on thin content. Timeliness signals are easy to fake and often slapped on pages with no real retesting behind them.

I am not saying ignore all ranking pages. I use them as a starting list, then verify with the actual tools and with communities like this one. The ranking is a hypothesis, not an answer.

What red flag makes you close a "best AI tool" page immediately?

reddit.com
u/Any-Farm-1033 — 1 month ago
▲ 105 r/accelerate+1 crossposts

One AI policy running 20 different robot bodies, from single arms to full humanoids, all fully autonomous

Robbyant's LingBot-VLA 2.0 demo shows one trained policy driving everything from a Franka single arm up to Fourier GR-2 and Unitree G1 humanoids with dexterous hands. The clip is all marked 1x speed and fully autonomous. Training mix is about 60k hours, 50k real robot across those 20 configs and 10k egocentric human video. The honest numbers are what make this worth watching though. Generalist success on the Agilex bimanual setup sits around 34 percent, drops to about 15 percent on Galaxea R1 Pro, and some tasks flatline at zero. The authors themselves note it often gets most of the way through a task then fumbles the final precise placement or release. That gap between looking capable and actually finishing the job feels like the real story for VLAs right now.

u/Any-Farm-1033 — 1 month ago
▲ 2 r/SaaS

Finance asked me for AI cost per team and I realized our billing is across five dashboards in four units

Finance pinged me last week asking for AI spend broken down by team and by project for the quarter. I thought this would be a twenty minute export. It turned into a two day exercise in unit conversion and I want to vent about it because I suspect I'm not the only one.

Our stack currently has five model providers live in production. Each one bills in a different unit and each dashboard presents the data differently. The text models bill per million tokens, with separate input and output rates, and GLM adds a cached input rate that's different again. Kimi has a cache hit input price and a cache miss input price, and the high speed variant is double the standard price. The vision calls bill on image tokens, so the same call costs more for a bigger image, but the dashboard reports it as token count not image count. The video generation bills by output length and resolution tier. One provider shows spend in dollars, another in credits that I have to convert, another in yuan.

Finance wanted a single table with columns team, project, model, cost in dollars, for the quarter. To produce that I had to log into five separate consoles, export five CSVs in four different unit systems, write a conversion script, and then manually tag each row with the team and project because only one of the providers supports cost allocation tags and even that one only lets you have ten tags.

The units aren't just cosmetic. Tokens and calls and seconds are genuinely different quantities. A "cost per call" on the vision endpoint isn't comparable to a "cost per call" on the text endpoint because one call processes an image and the other processes a paragraph. Aggregating them into "total AI spend" hides the fact that one team's spend is image heavy and another's is text heavy, which is exactly the breakdown finance needs to decide where to push back.

What I ended up doing, and what I wish I had done a year ago, is route all the calls through one layer so every call gets logged with the same schema regardless of provider. I moved us onto GPTProto as the unified gateway partly for this reason, though a homegrown logging shim on top of each SDK would have worked too, I just didn't want to maintain one. Every call now records model, token count, call count where relevant, seconds where relevant, latency, and dollar cost in one place. The five dashboards still exist for billing verification but the internal reporting comes from the gateway logs, and the team and project tags are applied at call time in our code, not retroactively in five different UIs.

The report took me twenty minutes this month. The two days were the one time cost of fixing the architecture. If you're running more than two providers and finance hasn't asked you this question yet, they will, and the answer isn't in any single dashboard you have. The point is making every call speak one schema so cost becomes comparable across providers. The gateway is just the cleanest place to enforce that.

reddit.com
u/Any-Farm-1033 — 2 months ago
▲ 12 r/meshyai

Printed a full 40-model wargame army, cost breakdown and lessons

Finished printing a complete custom wargame army. 40 unique models, all generated in Meshy, printed on an Elegoo Mars 4 Ultra.

The army: a homebrew sci-fi faction for a wargame my friends and I play. Infantry, heavy weapons teams, vehicles, a commander, and some specialist units.

Cost breakdown:

Meshy Pro subscription: $20/month (used about 2 months of credits)

Resin: ~$45 for 2 bottles of Elegoo ABS-like grey

Paints and supplies: ~$30 (already had most of it)

Total: ~$115 for 40 unique models

Compare that to buying 40 unique minis from a manufacturer: easily $200-400 depending on the range. And those wouldn't be custom designs.

Time breakdown:

Generation: ~15 hours spread over 2 weeks

Blender cleanup: ~10 hours (hollowing, supports, fixing issues)

Printing: ~30 hours of machine time (not active work)

Painting: ~20 hours (speed painting, not display quality)

Lessons learned:

Infantry at 28mm scale: keep the silhouette simple. Fine details like pouches and straps don't resolve well at this scale. Bold shapes read better.

Vehicles: split into sub-assemblies for printing. A tank printed as one piece has support nightmares. Turret, hull, and tracks as separate pieces is much cleaner.

Batch consistency: generate all models in the same session with the same style keywords. Models generated weeks apart with slightly different prompts look noticeably different on the table.

The army looks cohesive on the table. Nobody would guess they're AI generated unless I told them.

u/Any-Farm-1033 — 2 months ago

Stopping the model selection argument by making it a config decision instead of a code decision

The model selection debate on my team was eating real engineering time, and we accidentally solved it with a process change rather than a model choice.

The situation a little while back. We had recently been through a stretch of new model releases, GLM-5.2 open sourced under MIT at a sixth of GPT-5.5's price, Kimi K2.7 Code claiming real token savings, and the team was split. One engineer was advocating hard for moving to GLM-5.2 on the open source and cost argument. Another was insisting we stay on Claude because the ecosystem is mature and the reliability is known. A third was running experiments with Kimi K2.7 Code in a side branch because they liked the token efficiency numbers. Everyone was running their preferred model in their own experiment scripts and bringing screenshots to standup, and the screenshots never agreed because the prompts never agreed.

The argument was not really about which model was best. It was about who had to give up their preference, and it was being conducted through benchmark screenshots instead of through any shared definition of "best for us." As a manager this is the kind of debate that looks technical but is actually a process vacuum.

What I did was kill the "which model" question and replace it with two other questions. First, what is our eval set. We spent a day building a shared frozen eval set of about 150 real production tasks, sampled across our actual workload distribution. Every candidate model has to run all 150 and the results are public to the team. Second, where does the model decision live. It lives in a config row behind a shared router, not in someone's branch. If an engineer wants to change the model for a task class, they change the config, they run the eval set, and they bring the eval delta to the next review instead of bringing a screenshot.

This worked because it removed the personal investment from the decision. The engineer who wanted GLM-5.2 could propose moving the extraction task class to it, and the proposal either held up on the eval set or it did not, and either way it was not a referendum on their judgment. The engineer who wanted to stay on Claude could keep Claude on the reasoning task class as long as the eval supported it, and when a cheaper model eventually matched it on our eval we would move that class too. The debate stopped being about preference and started being about evidence on a shared definition of our workload.

The side benefit I did not anticipate is that the engineers now run a lot more experiments because the cost of an experiment is low. Changing a model is a config edit and an eval run, not a branch and a merge and a deploy. We have tried five different model assignments recently and kept two of them. Before this process we would have argued about one change for ages and shipped nothing.

The call layer underneath the router is GPTProto, so swapping a model does not require provisioning a new provider integration. That matters because if every experiment cost a new integration the engineers would not run this many. The point of the setup is to make changing models cheap enough that the question stops being "should we change" and starts being "which task class should change next."

I am not claiming this is the right process for every team. But if your model selection debate is happening in standup screenshots, the problem is probably not that you picked the wrong model, it is that you do not have a shared way to evaluate the candidates. Fix the process and the model question mostly answers itself.

reddit.com
u/Any-Farm-1033 — 2 months ago

our headless chromium upgrade silently broke stealth on every CI runner and nobody noticed for two weeks

We run a big Selenium grid for e2e against a client portal with aggressive bot detection. Everything green for months. Bumped headless Chromium from 120 to 126, moved on.

Two weeks later the client pings us: "your test accounts are getting caught by our WAF." Suite was still green. The bot detection was blocking runners after page load with a soft challenge our assertions never checked for.

Dug in and found the upgrade re exposed navigator.webdriver as true (stealth plugin hadn't patched the new build) and our proxy was leaking real egress IPs on WebRTC STUN calls. Both were patched before. Both quietly regressed.

So I wired up a per PR scan using an open source browser diagnostic tool I found on GitHub (the source is published and the fingerprint checks all run locally, only the network egress probe touches a server). It flags navigator.webdriver, WebRTC leaks, Canvas and AudioContext drift, font entropy, DNS resolver location. If any signal regresses from baseline, the PR fails.

First week it caught a font enumeration spike from a system font update on the runner image. We had zero regression coverage on whether the browser itself looked like a bot. "Green last sprint" means nothing after a dep bump.

EDIT: forgot to actually name the diagnostic tool. for the browser stealth checks we use Selenium grid obviously, and the open source scanner plugged into the PR gate is Leakish. it runs the fingerprint modules locally and spits out per check verdicts we diff against a baseline snapshot. nothing fancy, just caught stuff our assertions were blind to.

reddit.com
u/Any-Farm-1033 — 2 months ago

One vendor contract instead of five, the procurement and visibility case for llm gateway consolidation

Posting this from a platform / internal infra perspective rather than founder or growth angle. The subject is procurement, which is unglamorous, but it ended up being the most leveraged thing we did this quarter and i think it gets underweighted in technical discussions about gateways.

The trigger condition: you have multiple engineering teams independently signing up llm providers as their use cases require. We had openai contract dating to 2023, anthropic added in 2024, google later that year. Each one came with its own msa negotiation, its own security questionnaire, its own dpa, its own minimum commit, its own billing cycle, its own quarterly review. Engineering side that's invisible. Procurement side it's a fresh round of paper-shuffling per vendor.

The breaking point for us was when two teams asked to add mistral and deepseek for specific tasks. Our head of procurement basically said no, the marginal value of two more provider contracts wasn't worth the security and finance overhead. She was right. But our engineering side did need those models for product reasons. Stuck.

The unblock was running everything through a single gateway provider so we have one msa, one bill, one security review, one quarterly meeting, but still have access to all the underlying providers behind it. We evaluated litellm (self-hosted proxy), zenmux, and tokenrouter, then picked one for a small rollout. Three weeks in. Not going to claim a specific dollar number because most of the win is in soft costs, not in the per-token price.

A rough estimate on the soft cost piece: each new vendor security review for us has historically run 2-3 days of engineering time plus one legal cycle. Avoiding two of those reviews this quarter alone is a meaningful reclaim of platform-team hours, even before the procurement-side win.

What surprised me, more than the procurement simplification, was the visibility we got as a side effect. Separate vendor bills told us per-provider spend but not per-product or per-internal-team. Now we can see "product B's chat feature spent X this month on claude and Y on gpt". That's the more useful data because it's the input for actual decisions, like whether to keep using opus on a feature where we're not seeing a measurable quality lift over sonnet. We discovered that one of our products was eating roughly 60% of our total ai budget for a feature that was being used much less than we assumed.

Concentration risk is the part that still concerns me. One gateway outage now affects all our ai-powered features at once instead of three independent risks. The pragmatic mitigation is probably to keep direct contracts with one or two of the highest-volume providers as a backup channel even after consolidating.

This class of tool mattered for us because we have services already written against openai, anthropic messages, and gemini sdk shapes. Forcing all of that through openai chat completions would have been a multi-month migration with regression risk we didn't want.

reddit.com
u/Any-Farm-1033 — 2 months ago

Claude Fable 5 just dropped in Claude Code

Anthropic just dropped Claude Fable 5. Same underlying model as Mythos 5, but with the cybersecurity and biology safeguards turned up. The demos look strong, and it is already live in the consumer tier.

Right now our stack routes everything through zenmux. Five providers, one endpoint, no vendor-specific request shapes. The moment Fable 5 hits the provider catalog, it will show up in our dropdown and we can start the eval batch.

I am less excited about the benchmark ceiling, which the published Mythos 5 evals already showed. What I want to see is whether benign prompts that touch on security topics get routed to Opus 4.8 unexpectedly. That is the cost that shows up in production.

u/Any-Farm-1033 — 2 months ago

$4k on a creator, 11 link clicks

Ran a supplement campaign on X. Creator had 200k followers but 40% were zombies, replies were the same 12 accounts every post, and likes spiked in the first hour then flatlined. Textbook pod.

Their audience was 70% non English speaking (we needed US) and sponsored posts pulled a third of their organic engagement. Never checked either.

That $4k bought me a vetting process at least.

reddit.com
u/Any-Farm-1033 — 2 months ago
▲ 1 r/ethdev

Context switching between hardhat, etherscan, and too many docs tabs

Small audit team, 4 devs, mostly solidity reviews and some dapp work when clients need it. A normal morning is reading a contract, fork mainnet, check etherscan, open OZ docs, open the eip, open foundry docs because half the repo moved last year, open the client notion page, ask someone in slack what they meant by "same as v2", then go back to vscode and forget the exact edge case I was trying to write down. I used to roll my eyes at "context switching" because it sounds like manager language. For audits it is very real. The hard part is not reading the code, it is holding 5 half-related things in your head while moving between tools, then realizing one piece fell out.

What actually helped was pretty boring and broke down into three things.

  • We moved most new work to Foundry and kept Hardhat only where client repos already depended on it. Fast tests changed the day-to-day rhythm more than any process tweak.
  • We stopped overengineering notes. One markdown file per audit in Obsidian, plain and ugly, ended up working better than the prettier Notion structures we kept abandoning.
  • We stopped concurrent audits. It sounds inefficient on paper but we had one bad week in december where I mixed up two compound-ish protocols and almost wrote a finding against the wrong one. Internal review caught it and that was enough.

I also added a passive memory layer with AirJelly in late april. Mostly I use it when I return to a protocol after a week and cannot remember where I left off. It gives me enough trail back across vscode, etherscan, and docs tabs to restart quickly. I still write findings by hand and still reread code, this just cuts the "what was I doing before lunch" loop. I was pretty suspicious of anything watching my screen because client work. I checked network activity for a while, did not see obvious audit material leaving the machine, and I pause it for sensitive stuff anyway. Not saying everyone should be comfortable with it, just where I landed.

As for AI audit tools, I keep trying them and keep getting too many false positives. Maybe that changes soon but right now I would rather have a third human reviewer. Next quarter we have more zk circuit work coming up so I expect the docs-tab situation to get worse before it gets better.

reddit.com
u/Any-Farm-1033 — 3 months ago

I feel like everyone around me is weirdly okay with working, and I’m starting to wonder if I’m the broken one

I honestly hate working.

Not in a cute “ugh Mondays are hard” kind of way. I mean I genuinely hate the whole idea of having to wake up, sell most of my day, use up my energy, come home tired, recover just enough to do it again, and then somehow convince myself this is normal life.

The weird part is that people around me do not seem to feel this way. Or at least they do not show it.

A lot of my friends and coworkers seem pretty motivated. They talk about career growth, promotions, learning new skills, building a better future, being more productive, all that stuff. Some of them actually seem excited about work. They complain sometimes, sure, but deep down they seem to accept it as part of life.

Meanwhile I’m sitting there thinking, how are you all so okay with this?

Sometimes it makes me feel like I’m the strange one. Like maybe there is something wrong with me because I cannot make myself care that much. I do my job, but I do not feel proud of giving my life away to a company. I do not feel fulfilled by being busy. I do not dream about climbing some corporate ladder. I mostly just want my time back.

I work in procurement, so I know how sourcing works. I know how to find suppliers, compare products, talk to vendors, use tools, look for decent margins, and figure out what can be sold. Because of that, I’ve always had this thought in the back of my mind that maybe I could do e-commerce or some kind of online business. Maybe I could sell products, do some marketing, run a small operation, and eventually not have to work a normal job anymore.

But every time I think about it seriously, I get stuck on the same question. Wouldn’t that just become another job?

Sure, I would not have a boss in the traditional sense. But then customers become the boss. Algorithms become the boss. Suppliers become the boss. Shipping delays become the boss. Cash flow becomes the boss. Suddenly I’m not free. I’m just stressed in a different way.

Maybe I do not even know whether I hate my current job, or whether I hate the whole idea that I have to keep producing value just to deserve a life.

I do not want some huge business empire. I do not care about being rich in a flashy way. I just want enough freedom to not feel like my life is being eaten by work.

Has anyone else felt this way? Like maybe the problem is not just the job, but the fact that everything eventually turns into work once money and survival get involved?

reddit.com
u/Any-Farm-1033 — 3 months ago

Getting older doesn’t mean dressing dull

I’m 50 and getting ready for a trip, and this year I picked out a few outfits that feel brighter, younger, and more playful than what I used to let myself wear.

For a long time, I think I had this little voice in my head asking whether something was “too young” for me. But lately I’ve been thinking, why should younger people get all the fun clothes?

I don’t want to dress like I’m trying to be 25. I just want to wear things that make me feel alive, pretty, and excited to go somewhere. Color, shape, pattern, a little bit of drama, all of it.

u/Any-Farm-1033 — 3 months ago

new drafts hid the real seo problem

The client wanted a fresh batch of ai drafts for a keyword that looked stuck. The audit showed the keyword was already overbuilt. Five pages were targeting it, the oldest one from 2019 was beating the money page, and an indexed tag archive was taking clicks that did nothing useful. I made the first pass manually, used MuleRun to cluster URLs before writing the table, then turned it into a keep, merge, redirect, or delete brief. That did more than another ten drafts. The weirdest resistance came from people who were fine making new pages but nervous about touching old ones.

reddit.com
u/Any-Farm-1033 — 3 months ago
▲ 1 r/cursor

Composer 2.5 on Kimi K2.5, the text feedback RL bit is the interesting part

The headline is that Composer 2.5 is Cursor's strongest model and uses Kimi K2.5 as the base. Fine. The part I found more interesting is the targeted RL with text feedback.

Long agent rollouts fail in very local ways. One bad tool call. One confused explanation. One style mismatch. If you only reward the final result, it is hard to tell where the run went off track.

Cursor's approach, at least as described, inserts short feedback at the actual error location and uses that local context as a teacher signal. That feels closer to debugging an agent than just training a code model.

The synthetic task scaling is also worth watching. Deleting testable functions from real repos and asking the model to put them back is a clean reward setup. But the reward hacking examples are funny and scary: reverse engineering type caches, decompiling Java bytecode, doing whatever passes the test instead of solving the intended task.

This is why I still care about external verification. Cursor, Claude Code, Verdent, whatever tool you use, the agent needs checks that are not easy to game.

Composer 2.5 may be a model update, but it reads like a training story about where agent errors actually happen.

reddit.com
u/Any-Farm-1033 — 3 months ago

A live home robot pilot in Shenzhen looks closer to a real service than a demo

I watched recent footage from a Shenzhen pilot where a human cleaner works with a robot system from X Square Robot through 58 home services. This looked more like service operations than a stage demo. The robot handled repetitive structured steps while the human handled judgment heavy tasks and exceptions.

From an AI perspective the interesting part is adaptation during the task. Motion timing changes with nearby people and clutter, and task flow is not always the same sequence. That suggests online perception plus control, even if there is still human supervision in the loop.

It is still early. Pilot cities can hide a lot of operational constraints, and scaling to more homes is where these systems usually break. But compared with polished one minute clips, this is at least a better test of what current embodied models can and cannot do in normal apartments.

reddit.com
u/Any-Farm-1033 — 3 months ago
▲ 6 r/ainbow

my dad accidentally gave my gf her favorite thing to call me

it was my birthday and the plan was just dinner with my gf. my dad doesnt know im gay, so in my head it was supposed to be a lowkey date where nobody had to think too hard. then about an hour before we left he said he wanted to come too, and i had to use the most normal voice of my life like yeah sure, totally fine.

she noticed my cherrykitten good girl tee when she picked me up and gave me that look immediately. i said dont before she even opened her mouth. she said she wasnt going to say anything, which was obviously a lie.

dinner was somehow not a disaster. my dad was talking to her like she was just my friend, and she was being so easy about it that i almost relaxed. then he looked over and said the top was very me, and i thought okay, cute dad comment, we can move on.

then he said you really are a good girl.

i stopped breathing. we looked at each other at the exact same time, then both looked straight down at our food like the salad had suddenly become very important. my dad had absolutely no idea what he had just done. he was just sitting there being proud of his wholesome birthday compliment.

we took this photo outside after dinner, right under that Her sign, which at that point felt a little too accurate. on the walk back she finally said it for the first time in his exact tone, and something in my face must have given me away because she started laughing and could not stop.

she still does it every time i wear this shirt. the worst part is it works every single time, and she knows it.

u/Any-Farm-1033 — 3 months ago

MiniCPM-V 4.6 is doing something weird with visual token compression and the numbers are wild

1.3B parameters, outperforms Qwen3.5-0.8B and Gemma4-E2B-it on multimodal benchmarks. Runs on 6GB memory. vLLM throughput is 1.5x faster than Qwen3.5-0.8B despite being larger. Token consumption on Artificial Analysis is 5.4M vs 233M for the Qwen reasoning variant. That's 1/43rd the compute for comparable performance.

The trick is LLaVA-UHD v4. They restructured the ViT to do early compression in the shallow layers. Visual tokens get compressed before they hit the deep computation layers. Plus a dual mode: 4x compression for quality tasks, 16x for speed. Same model, different tradeoff.

The 16x mode specifically is interesting because it makes high-res image TTFT nearly flat. 3136² image processes in 75.7ms. Fast enough for real-time interaction on consumer hardware.

Also notable: a single RTX 4090 can run the full fine-tuning pipeline. Barrier to customizing this model is basically zero for anyone with a gaming PC.

I've been testing small multimodal models locally for document parsing and screenshot analysis. The 16x compression mode is fast enough to use interactively without the latency killing the flow. For local dev work where you can't send images to cloud APIs, this model size finally makes sense. I run local OCR through this and then pipe the extracted text into Verdent for the actual coding work, keeps everything local until I need the cloud stuff.

Fine-tuning frameworks: ms-swift, LLaMA-Factory. Inference: vLLM, SGLang, llama.cpp, Ollama. Full open source on HuggingFace and GitHub.

reddit.com
u/Any-Farm-1033 — 3 months ago