r/gptimage2prompts

pagedMark: invisible SynthID-class watermark removal for AI images (ChatGPT, gpt-image, DALL·E, Sora, Gemini, Nano Banana), running on Metal
▲ 9 r/gptimage2prompts+5 crossposts

pagedMark: invisible SynthID-class watermark removal for AI images (ChatGPT, gpt-image, DALL·E, Sora, Gemini, Nano Banana), running on Metal

pagedMark removes AI provenance from content you generated yourself. Two different things, and it is worth separating them. The first is metadata: C2PA Content Credentials, EXIF, XMP, IPTC, the generator parameters. That part is easy and verifiable, and a screenshot does it too. The second is the invisible pixel watermark that a screenshot does not touch, the SynthID class of marks, which has to be disrupted by regenerating the image itself.

Coverage on the image side is ChatGPT, gpt-image, DALL·E, Sora, Gemini and Nano Banana for the invisible marks, plus a registry of visible vendor labels (Doubao, Jimeng, Qwen, Kling, Yuanbao, Baidu, LibLibAI, Samsung Galaxy AI). On the video side it handles the visible marks from Sora, Veo, Seedance, Dola, Hailuo and Kling, and the metadata that travels with them.

The reason this is worth a post rather than a link is that it is built for Apple Silicon instead of ported to it. I spent several days getting the pipeline to run correctly on an M5 with 16 GB, meaning predictable and measured rather than merely launching. Most of what I assumed turned out to be wrong, so the measurements are below.

The four-step distillation LoRA invents texture, and more steps make it worse

A low strength edit runs the tail of a long schedule: strength 0.15 executes the last four steps of twenty seven. A LoRA distilled for four timesteps spanning the entire noise range is off its distribution there. Wherever nothing conditions the model, and flat dark fabric gives a Canny ControlNet no edges at all, it fills the gap from its prior. On a night photograph that arrives as coloured camouflage across black clothing.

Global stage, 1448x1080, strength 0.15, seed 0 Invented texture PSNR Wall
Lightning, 4 steps 1.73x source 28.54 dB 41 s
Lightning, 8 steps 1.80x 28.19 dB 29 s
Lightning, 16 steps 1.84x 27.85 dB 62 s
Undistilled base, 16 steps 1.19x 29.25 dB 71 s
Undistilled base, 24 steps 1.20x 29.17 dB 132 s

Asking the distilled model for more steps made the artifact worse, which is what identified the distillation rather than the step count as the cause. Dropping the LoRA costs roughly three times the wall time and buys back both fidelity and correctness.

Three wrong theories I paid for first, in case they save someone else the time. Not the fp16 VAE: a bare encode and decode round trip of the same crop is clean in fp16 and in fp32, tiled or whole, at 34.6 dB. Not Metal's fp16 in general: bf16 measured marginally worse. Not Canny picking up sensor noise: the Canny map of that region is completely empty, which was the actual clue.

Metal pages instead of failing, so memory has to be measured

torch.mps.recommended_max_memory() reports 11.84 GiB on a 16 GB machine. Exceed it and nothing raises an error. The process starts swapping, and a run that should take 23 seconds takes an hour instead.

With VAE tiling disabled, a 1.57 MP frame peaks at 18.74 GiB and takes 59 seconds. With tiling it peaks at 10.92 GiB and takes 23 seconds. So tiling carries real weight on a small machine, but its boundaries leave a faint texture, which is why it is now decided per frame from the device budget rather than switched on globally.

Diffusion untiled at 2.5 MP went into swap and did not finish within twelve minutes. Tiled at 1024 px, a 5.07 MP frame holds 10.93 GiB, finishes in 88 seconds, and keeps its native geometry.

Sequential CPU offload works on MPS, and it is what makes 8 GB usable

The stack is 7.7 GiB of weights. An 8 GB Mac reports a working set of roughly 5.3 GiB, so it does not fit however the activations are handled. Streaming the weights one module at a time:

Same frame, same seed Peak device memory Wall
Weights resident 7.70 GiB 7.1 s
enable_sequential_cpu_offload(device="mps") 0.28 GiB 24.1 s

Twenty seven times less peak memory for 3.4 times the wall time. The plan is chosen from the measured budget and then printed, because a run three times slower than the fast path looks broken unless it says why.

Two Metal gaps worth knowing if you are porting anything

torch.float8_e4m3fn does not exist on MPS at all. The error is RuntimeError: Undefined type Float8_e4m3fn. Any pipeline that streams float8 weights, which several VRAM managed stacks do, cannot be loaded there under any configuration.

SAM's processor emits its box and point prompts as float64, which Metal also has no type for, so moving the batch to the device raises rather than degrading. A single cast fixes it, but nothing tells you that is the problem.

The expensive one: fp16 sampling on MPS returns zeros silently

I added a memory optimisation that encodes the two fixed prompts once and drops the text encoders, saving a measured 1.52 GiB of the 8.79 GiB the loaded stack holds. Two of four face crops then came back as all zero black rectangles. Deterministically, at the same seed, with nothing raised anywhere.

The embeddings were innocent. CPU fp16, MPS fp16 and fp32 encodings of that prompt agree to 0.0009 on tensors with a standard deviation of 3.06, and the same crop generated in isolation is correct either way. Freeing unrelated memory changed the allocation pattern the crops met after the global pass, and that alone was enough. I withdrew the optimisation and added a guard that drops any empty crop instead of compositing it.

If you run fp16 diffusion on Metal, check your output for degeneracy. It will not tell you.

What it does not claim

Regeneration is not payload deletion. The image changes: faces, text and fine detail move, and the numbers above are the measured size of that change rather than a reassurance.

No public local decoder exists for SynthID class marks, so identify reports unknown and never clean. Verification is the provider's verifier or nothing. The 0.15 operating point comes from the upstream project's record against openai.com/verify on CUDA. I have not re-run that check on Metal, and Metal is not bit identical to CUDA, so I am claiming the same operating point and not the same verdict.

It is for content you generated or own. The visible mark registry accepts AI generation labels only. Stock agency previews, marketplace and classifieds watermarks are deliberately out of scope, and that boundary is in the repository rather than only in this comment.

Because "how much did that cost my picture" is the whole question

pagedmark measure before.png after.png

PSNR over the frame, PSNR per detected face, and how much mid band structure appeared where the source was flat and dark. The third metric is the one that caught the camouflage, and it took two attempts. Per pixel chroma statistics rank the artifact below the source, because the source's own sensor grain carries more per pixel variance than the invented blotches do. A plain band ratio fails too, since any linear filter reports doubled grain and doubled blotches identically. Normalising mid band energy by fine detail energy, against the same ratio in the source, measures the shape of the spectrum instead of its size.

uv tool install "pagedmark[diffusion]"
pagedmark invisible photo.png -o clean.png
pagedmark invisible photo.png --preview      # 46.6 s instead of 112.6 s

Code: https://github.com/doofzoff/pagedMark
PyPI: https://pypi.org/project/pagedmark/

Happy to answer anything about the Metal specifics. That is the part I would have wanted written down before I started.

u/d0ofz — 11 hours ago
▲ 6 r/gptimage2prompts+1 crossposts

Hey everyone! I'm u/ElasticAIGirl, a founding moderator of r/gptimage2prompts.

This is our new home for all things related to {{ADD WHAT YOUR SUBREDDIT IS ABOUT HERE}}. We're excited to have you join us!

What to Post
Post anything that you think the community would find interesting, helpful, or inspiring. Feel free to share your thoughts, photos, or questions about {{ADD SOME EXAMPLES OF WHAT YOU WANT PEOPLE IN THE COMMUNITY TO POST}}.

Community Vibe
We're all about being friendly, constructive, and inclusive. Let's build a space where everyone feels comfortable sharing and connecting.

How to Get Started

  1. Introduce yourself in the comments below.
  2. Post something today! Even a simple question can spark a great conversation.
  3. If you know someone who would love this community, invite them to join.
  4. Interested in helping out? We're always looking for new moderators, so feel free to reach out to me to apply.

Thanks for being part of the very first wave. Together, let's make r/gptimage2prompts amazing.

reddit.com
u/Bubbly_Total6443 — 2 days ago
▲ 0 r/gptimage2prompts+1 crossposts

GPT Image 2 + Seedance 2.5 made this lost wallet moment feel weirdly real

We tried a simple everyday storytelling test with GPT Image 2 + Seedance 2.5.

A woman drops her wallet on a busy city sidewalk, a stranger notices, picks it up, and returns it before she disappears into the crowd.

What I liked about this one is that the idea is super small and human — no explosions, no crazy concept, just a believable “good deed in the city” moment. That actually makes it a better test for realism, acting, and natural background behavior.

The parts I wanted to push were:

  • authentic handheld documentary feel
  • believable city movement and pedestrian flow
  • natural reactions instead of overacting
  • subtle emotional storytelling
  • everyday realism from start to finish

The best part is the reaction shot when she realizes she lost the wallet and genuinely looks relieved. Feels like the kind of moment you could randomly capture in real life.

Prompt:

"Scene 1 — Busy Morning | 0:00–0:05
“Photorealistic handheld street footage of a young Asian woman in her mid-20s walking through a busy city sidewalk in the morning. She wears a beige jacket, white shirt, loose blue jeans and carries a small shoulder bag. Natural pedestrians, bicycles, buses and storefronts move around her. She checks her phone while walking, completely unaware of what is about to happen. Authentic documentary-style video.”

Scene 2 — The Discovery | 0:05–0:10
“Low handheld camera angle following behind the woman as she walks away. A black leather wallet accidentally slips from the side pocket of her shoulder bag and lands on the sidewalk. She continues walking without noticing. A young man walking several steps behind suddenly notices the wallet on the ground.”

Scene 3 — The Decision | 0:10–0:15
“Medium handheld shot of the young man picking up the wallet and looking around for its owner. He opens it briefly only enough to identify who it belongs to, then immediately closes it. He spots the woman disappearing into the crowd and starts walking quickly after her. Realistic expressions and natural body movement.”

Scene 4 — Catching Her | 0:15–0:20
“Dynamic handheld tracking shot of the young man moving through the busy sidewalk, carefully avoiding pedestrians as he tries to catch up with the woman. He finally reaches her and gently taps her shoulder. She turns around with a confused expression, then looks surprised when he holds out her wallet.”

Scene 5 — The Reaction | 0:20–0:25
“Close realistic shot of the woman taking the wallet back and checking her bag in disbelief. She realizes she never noticed losing it. She looks at the man with genuine gratitude and smiles. He simply smiles back and gestures that everything is fine. Natural emotional acting, no exaggerated expressions.”

Scene 6 — Small Good Deed | 0:25–0:30
“Wide handheld shot as the two people walk in opposite directions through the busy city street. The woman briefly looks back and smiles. The man continues walking normally as the city moves around him. Warm late-morning sunlight, realistic pedestrians and traffic, subtle camera movement, authentic everyday-life atmosphere.”

Overall style:
“Ultra-realistic live-action footage, documentary cinematography, natural human behavior, realistic skin and clothing, authentic city environment, imperfect handheld movement, believable background activity, natural lighting, subtle motion blur, no cinematic fantasy effects, no logos, no text, no watermark.”"

Share your thoughts in the comment section below!

u/FarReputationAI — 6 days ago

Need a Prompt for Business

Hey , new in the community , looking for a prompt , which incan create realistic models for my business, which I can use for posters , and ads . Please helpp 🙏🏻

reddit.com
u/AYVIIIIIII — 8 days ago

First attempt at creating an AI influencer

Don't claim to be an expert, strictly an amateur who is looking to learn more about AI and improve. I like the way this prompt/image came out. Prompt below.

Use the uploaded reference image ONLY to preserve the adult woman's facial identity, facial structure, skin tone, eye color, and general body proportions. Do not copy its pose, clothing, background, lighting, or expression. Create the following scene entirely from this written description.

Create an ultra-photorealistic vertical office portrait of a glamorous adult blonde woman standing beside a desk in a modest professional office. The result must look like a real high-end DSLR photograph, not an illustration, CGI render, or composite.

COMPOSITION:

Vertical portrait, roughly 9:16. Frame from slightly above the top of her hair to mid-thigh. She occupies about two-thirds of the frame and stands slightly left of center. Keep enough office visible to establish the setting.

A dark brown office desk enters diagonally from the lower-right corner toward her. Near the camera is a thick stack of white printed paperwork with many layered page edges, slightly softer in focus than the woman.

CAMERA:

Use a full-frame 75–85 mm portrait-lens look. Camera several feet away at approximately upper-torso/chest height, almost level. No wide-angle distortion. Natural face, arm, torso, and body proportions. Moderate shallow depth of field, approximately f/2.8–f/3.5: eyes, face, nearby hair, upper body, and clothing crisp; hands nearly sharp; background and foreground paperwork gently softer but recognizable.

POSE:

Her pelvis and torso are turned about 35–45 degrees away from the camera toward frame-right, creating a graceful three-quarter side view. She is not square to the lens. One hip shifts subtly backward, creating a natural S-curve through the waist without exaggeration.

She leans only slightly forward toward the desk, about 5–10 degrees, not dramatically bent. Her back remains naturally straight.

Both arms extend downward toward the desk at frame-right. Elbows are almost straight but relaxed. Both hands lightly rest against the desk near its edge, one slightly farther forward. Fingers are relaxed and naturally separated. Her weight is supported mainly by her legs, not her arms.

Her head turns back over her shoulder toward the camera while her torso stays angled away. Chin is level or slightly lowered. The body and face create clear counter-rotation: hips and chest angled to frame-right, face turned back toward frame-left.

EXPRESSION:

She does not look directly into the lens. Her eyes glance sideways toward frame-left, as if noticing someone nearby. Expression is calm, confident, subtly playful, and mildly amused. Lips are closed or barely parted with a tiny asymmetric smile. No broad grin, teeth, exaggerated seduction, surprised eyes, or dramatic eyebrow raise.

FACE:

Preserve the identity from the uploaded reference. Render her naturally at this new angle rather than pasting or transplanting the face. Maintain realistic adult facial anatomy, natural cheek volume, jawline, nose, lips, eyelids, and eye spacing.

Skin is warm and lightly sun-kissed with realistic pores, tonal variation, and restrained highlights. Makeup is polished but believable: warm neutral foundation, subtle bronzer, peach blush, neutral brown eye makeup, fine eyeliner, mascara, and muted rosy-nude lips. Avoid waxy, porcelain, plastic, or over-retouched skin.

HAIR — VERY IMPORTANT:

Extremely long warm honey-blonde hair reaching approximately the lower back, with golden blonde, honey blonde, lighter sun-kissed strands, and slightly darker blonde roots.

Across the top and crown are multiple narrow interwoven braids. Several small braids begin near the front hairline and temples, sweep backward, overlap, and wrap around the crown like a loose multi-braid crown. They must look individually woven from real strands with visible three-dimensional texture. Do not make one thick rope braid or rigid halo.

Most hair remains loose below the braided crown. It is very long, thick, voluminous, and finely crimped/wavy, with narrow irregular waves rather than large salon curls. It falls behind her shoulders and down her back in many visible strands. Include soft face-framing pieces near the temples and cheeks and fine flyaways along illuminated edges. Hair obeys gravity.

TOP:

She wears a fitted muted chocolate-brown/taupe-brown long-sleeve knit top made from fine matte stretch jersey, close-fitting through the waist and torso.

The top has pronounced cold-shoulder cutouts exposing both shoulder caps while the sleeves reconnect below and continue to the wrists. The neckline contains a deep narrow V-shaped decorative opening/panel extending down the upper chest. Its border is decorated with small silver or clear reflective sequins, tiny mirrored pieces, or metallic beads catching isolated highlights. Include small brown fabric bridges/straps crossing parts of the V.

Show realistic knit texture, seams, mild tension folds at the waist, gathering around elbows, and natural forearm wrinkles.

SKIRT:

Very short light beige, cream, or pale-stone pleated mini skirt. Structured, moderately heavy opaque fabric. Broad architectural box pleats and knife-pleat-like folds with large overlapping panels and slight asymmetry around the hips. It sits around the low waist/upper hips and flares gently outward. Hem reaches upper thigh. Pleats are broad and sculptural, not dozens of tiny school-uniform pleats. Each fold follows gravity and casts soft shadows.

BODY:

Maintain the general figure and proportions of the adult woman in the uploaded identity reference. Slender, feminine, naturally proportioned, with a defined waist, proportionate shoulders and ribcage, natural hip width, realistic arms, torso, and thighs. Let the pose create the silhouette. No extreme hourglass, implausibly tiny waist, enlarged chest, oversized hips, stretched legs, or distorted spine.

OFFICE:

The office is practical and ordinary, not luxurious. Walls are warm cream or beige.

At frame-left behind her is a tall black shelving/storage unit. The upper shelf contains several upright navy-blue ring binders with pale labels and a few lighter binders. Lower shelves contain stacked folders, loose papers, and ordinary office supplies. Beneath or beside it are matte black filing drawers/cabinets with simple horizontal silver handles.

At frame-right is a cream wall with two or three black wall-mounted document/file holders containing white sheets. Keep writing small and indistinct.

Along the lower-right background is a low black cabinet or credenza. Several upright books stand along its top with muted blue, tan, burgundy, black, and neutral spines.

The foreground desk is dark brown polished wood. A large stack of white paperwork sits close to the camera at lower-right with realistic page edges and faint printed lines/forms.

LIGHTING:

Soft believable daytime office light. Primary light comes from camera-right/front-right like a large unseen window. It creates warm highlights on her face, forehead, cheekbone, exposed shoulder, hair, and forearm while the opposite side remains softly filled.

Neutral-warm white balance around 4800–5200K. Hair has delicate warm edge highlights revealing individual strands. Skin highlights are soft and natural. No neon, hard flash, orange-and-teal grading, dramatic spotlight, heavy rim light, or blown highlights.

PHOTOGRAPHIC RENDERING:

High-end full-frame DSLR/mirrorless realism. Detailed hair fibers, pores, eyelashes, knit fabric, sequins, skirt weave, paper edges, books, cabinets, and desk grain. Natural microcontrast, realistic dynamic range, fine photographic grain. Avoid HDR halos, oversharpening, fake bokeh, painterly texture, CGI sheen, or beauty-filter skin.

SPATIAL AND ANATOMICAL REALISM:

Everything occupies one coherent space. Hands physically meet the desktop. Five correct fingers per hand, realistic joints, nails, wrists, and spacing; no fused or duplicated fingers. Arms connect correctly beneath sleeves. Cold-shoulder openings reveal actual shoulders. Hair passes correctly behind and around shoulders. Skirt emerges naturally below the top. Background furniture follows consistent perspective and scale. Use one unified light direction, coherent depth of field, believable contact shadows, and correct occlusion.

OVERALL RESULT:

A polished candid moment in an ordinary office: an elaborately styled blonde woman stands beside a desk, body turned away in three-quarter profile, both hands resting lightly on the desktop, then turns her head back over her shoulder and gives someone outside frame a tiny knowing sideways smile. The contrast between glamorous styling and mundane office surroundings should feel striking but believable.

AVOID:

Frontal posing, direct eye contact, toothy smile, short or straight hair, missing braids, black top, plaid skirt, transparent fabric, other people, hats, glasses, prominent jewelry, tattoos, logos, watermarks, captions, borders, graphics, malformed hands, extra fingers, warped limbs, floating hair, distorted anatomy, exaggerated body proportions, excessive background blur, or a luxury-office redesign.

u/NeighborhoodDue4189 — 10 days ago