Which models Reliably use the Github Connector?

Since kimi k3 came out, I was thinking if there is one AI subscription that kimi potentially replaces, it's either perplexity or google. In contrast kimi will not reduce my subscription levels for openAI or anthropic. The reason being, is how I use kimi currently, I only use kimi for code reviews and to explain the answers from claude or gpt. This is because I don't have the trust history with kimi as I do with either claude or gpt. I know kimi is a great model and I probably don't have a reason to distrust less than any other model.

So here's the thing, when I use the kimi model on their website with the github plugin, it reliably searches github for me to answer questions about my project. When I do the same thing with perplexity sometimes the answers imply that it didn't use the github connector even when it was selected. Recently I noticed this when selecting the sonnet 5 model. It's possible that the connector was down (e.g. API limit exceeded), which might not be in perplexities control but kimi will work around this by using http to read raw github user content.

Recently, I used peprlexity and I selected "best" for model, and the answers implied it knew many things about my project. This might be a memories feature, but maybe also the github connector would have worked if I used it. I'm wondering if selecting "best" is the most reliable way for the github connector to work.

Also, perplexity has the kimi model, so in theory I can just use the kimi model on perplexity, but at the very least perplexity should tell me when the github connector fails, because this is essential for code reviews.

reddit.com
u/s243a — 12 days ago

Has anyone tried TestCircle?

I got an email about this testCircle app, and I was wondering if anyone tried it. One thing I like is you don't get to pick what apps you test, and this makes me think about what personal information an could collect, granted I could make that up.

reddit.com
u/s243a — 13 days ago

Has anyone tried TestCircle?

I got an email about this testCircle app, and I was wondering if anyone tried it. One thing I like is you don't get to pick what apps you test, and this makes me think about what personal information an could collect, granted I could make that up.

reddit.com
u/s243a — 13 days ago

[Android testers wanted] sciREPL — a multilingual notebook for R, Python, Prolog and more

I’m looking for testers for the free and Pro versions of sciREPL.

Picture something like Excel with one column, except cells contain code instead of formulas. Unlike Jupyter or Colab, different cells can use different programming languages within the same workbook.

sciREPL currently supports:

  • R
  • Python
  • Prolog
  • TypeR
  • JavaScript
  • Lua
  • Bash
  • ClojureScript

The demo video shows a “Breakfast Democracy” workbook. R loads and visualizes ranked-preference data, Python clusters and aggregates the rankings, and Prolog queries the resulting majority graph and identifies preference cycles.

The free version is open source and runs both as an Android app and in a browser. The Android-only Pro version adds an AI assistant that can connect directly to model APIs or to remote coding agents such as Claude Code and Codex.

Workbook data is stored locally on the device. The Pro version is normally paid, but closed-test participants will receive a 100%-off promotional code.

Google requires me to have at least 12 testers remain enrolled for 14 days before I can apply for public distribution.

How to participate

First join the tester group:

https://groups.google.com/g/scirepl-android-testers

Then opt into one or both tests:

SciREPL Free

https://play.google.com/apps/testing/com.unifyweaver.scirepl

SciREPL Pro

https://play.google.com/apps/testing/com.unifyweaver.scirepl.pro

Pro testers can contact me for their promotional code.

You can also try the free browser version immediately:

https://s243a.github.io/SciREPL/

Testing can be simple: install the app, open Browse Packages, Bundles & Workbooks, load a workbook, select Run all cells, and report any errors, confusing behaviour, or incorrect output.

Browser testing is welcome, but only Android Play Store enrolment counts toward Google’s closed-testing requirement.

u/s243a — 14 days ago
▲ 2 r/modelcontextprotocol+1 crossposts

I open-sourced SciREPL-MCP: connect an Android notebook to MCP clients, coding agents, and an optional remote shell

I’ve open-sourced the host-side MCP components of SciREPL under the MIT licence:

https://github.com/s243a/SciREPL-MCP

The main reason is trust and transparency. People can inspect the network-facing code and run the broker on a computer they control.

The broker exposes SciREPL’s approved notebook tools to MCP clients. It can also optionally connect SciREPL Pro to a coding agent such as Claude Code or Codex, a host shell, or both. Agent and terminal access are disabled by default, pairing-token protected, and intended to be reached through Tailscale Serve or SSH rather than exposed directly to the internet.

SciREPL is similar to Jupyter or Colab, except cells in the same workbook can use different programming languages. The free version is open source and runs as a PWA and Android app. The Pro Android client, including its AI panel and remote-bridge interface, remains closed source.

I’m also looking for Android testers. Use the same Google account for both steps.

First, join the tester group:

https://groups.google.com/g/scirepl-android-testers

Then opt in to either or both tests:

SciREPL Free:
https://play.google.com/apps/testing/com.unifyweaver.scirepl

SciREPL Pro:
https://play.google.com/apps/testing/com.unifyweaver.scirepl.pro

Pro testers can receive a separate 100%-off promotional code. I’d particularly appreciate feedback on the setup instructions, security model, and remote-agent experience.

reddit.com
u/s243a — 16 days ago
▲ 1 r/kimi

I asked Kimi K3 to design a cooling-tower water-recovery system to help reduce water use in data centres.

I was trying to think of a demonstration for sciREPL Pro and wondered: what if we could recover some of the water lost from cooling towers?

sciREPL is a Jupyter Notebook/Colab alternative. You can think of it as being a little like Excel, except it has one column and uses code instead of spreadsheet formulas. Tools like these are popular in data science.

The Pro version of my app has an agent panel that can connect directly to an LLM API or, alternatively, to a remote coding agent such as Claude Code or Codex.

GPT-5.6 Sol prepared the prompts given to Kimi for this demonstration. The first prompt asked Kimi to design the vapour-recovery system. The second prompt asked it to model a dry-cooling system for comparison.

The slides were also created by GPT-5.6 Sol, while the video editing and narration were done by me. The video still needs some work, but I spent several hours on it over two days—and I don’t have a production team!

The free version of the app is available on GitHub:

https://github.com/s243a/SciREPL

It does not contain the AI-agent panel included with the Pro version. I need testers before either app can appear on the Google Play Store. Testers will receive early access to the Pro version.

u/s243a — 17 days ago

Kimi K3 Is Impressive, but "Better and Much Cheaper" Is Too Simplistic

Kimi K3 is getting a lot of hype. Some claims say it beats Fable 5, GPT-5.6 Sol, even Opus 5. I don't buy the strong version. My read: Kimi K3 sits between the previous frontier tier (Opus 4.8 / GPT-5.5) and the current one (Fable 5 / GPT-5.6 Sol), genuinely good, but not quite there. On Artificial Analysis's Intelligence Index, Kimi scores 57, behind both Fable 5 and GPT-5.6 Sol, roughly level with Opus 4.8 and GPT-5.5. x

The benchmark headline problem

"Kimi beats Fable at X" often hides which X: frontend generation, a specific harness, an effort setting, or pass@k with multiple attempts allowed. DeepSWE shows this clearly, and the cost evidence here is genuinely mixed.

In one Kimi K3 Max vs GPT-5.6 Sol Max comparison, Sol wins pass@1 (72.7% vs 68.5%), but Kimi is cheaper per rollout ($4.65 vs $8.37) and pulls ahead at higher pass@k. A separate small programming micro-benchmark found Sol cheaper per correct answer than Kimi — but that wasn't DeepSWE, so it shouldn't be generalized. These aren't necessarily contradictory; they measure different things: one high-confidence attempt vs several cheap ones, cost-per-rollout vs cost-per-correct-solve. Anyone citing a single DeepSWE cost number without specifying which is skipping the part that matters. linkedin

Why I still rank it below

Interesting programming pulls from math, algorithms, systems tradeoffs, and domain knowledge outside the codebase. That's why broader reasoning benchmarks matter even for coding. They're a proxy for whether a model can transfer concepts when a task isn't "edit this function" but "figure out the right approach first."

The gap here is concrete. Fable 5 scored 88% on FrontierMath Tier 4, about 13 points above GPT-5.5's ~75%. Artificial Analysis also has Fable 5 leading its AA-Omniscience knowledge benchmark. GPT-5.6 Sol trails Fable by roughly a point on the aggregate Intelligence Index while costing about a third as much, and it topped GeneBench-Pro, a hard genomics/quantitative-biology benchmark, at 31.5% — a decent proxy for general scientific reasoning, if not coding directly. aiweekly

Kimi K3 doesn't show up as a contender on any of these. Its strengths sit in a different lane: frontend generation, some agentic coding, not the deep cross-domain reasoning the newest tier is winning on. That's the real basis for ranking it below Fable 5 and GPT-5.6 Sol: not just index position, but a measured gap in the cross-disciplinary reasoning that separates "good coding agent" from "frontier model."

API price ≠ task price

Kimi's tokens are cheap ($3/$15 per million vs Sol's $5/$30). But cheaper tokens don't guarantee cheaper tasks — longer runs, more turns, more retries eat the margin. Artificial Analysis found Kimi and Sol nearly tied on cost per task ($0.94 vs $1.04), despite the sticker-price gap. My guess: Kimi's edge holds on short, easy, cache-friendly work, and shrinks as tasks get harder. myclaw

Subscriptions are murkier still

I burned 6.87% of my monthly Moderato quota in a few hours doing GitHub-connected code review. That's not a controlled benchmark. It's one real data point missing from the hype videos. Kimi's docs confirm Agent, Deep Research, Kimi Code, and connectors all draw from one shared credit pool metered by token use. A $19/month price tells you little about how far that actually goes in real agentic work. kimi

One aside: engineer vs. scientist

Subjectively, Claude tends to commit to a complete implementation in one pass; GPT/Codex explores well but often needs more "continue" prompts to finish. That changes effective cost because finishing in one shot beats needing three follow-ups, even at a higher sticker price.

Bottom line

Kimi K3 is a legitimately strong near-frontier model, likely the better economic choice for easy-to-medium tasks. But "clearly better than Fable/Sol" and "obviously much cheaper" both overstate the evidence. DeepSWE cost comparisons point in different directions depending on setup — that's the actual state of the data, not a gap in this analysis. What would change my mind: a larger, harness-controlled study measuring cost-per-correct-completion across a real mix of easy and hard tasks.

reddit.com
u/s243a — 26 days ago

[sciREPL v1.0.1] Substantially improved TypeR support — testers wanted for Play Store launch

I'm releasing sciREPL v1.0.1, which brings major improvements to TypeR integration. sciREPL is a Jupyter-style notebook app that lets you mix multiple languages in a single worksheet.

**Why we forked TypeR**

The standard TypeR type system couldn't express variable-argument (variadic) inputs in its standard-library declarations. We extended it to properly type variadic R functions like `cat`, `paste`, `sprintf`, and `c`.

**What's new in v1.0.1**

* **Direct TypeR notebook cells** — no longer need to wrap most code in R blocks * **Named and heterogeneous variadic arguments**, including forwarding collected args into another variadic call * **Named** `#!source` **cells** for storing reusable TypeR source without executing it or producing output * **Prolog → TypeR compilation**, with generated code placed in a named cell ready for execution * **Two included typR workbooks**: "TypeR Introduction" and "Prolog Generates TypeR" — both accessible via *Browse Packages, Bundles & Workbooks* in the sciREPL menu, along with workbooks for other supported languages

**Other supported languages**

TypeR is one of several kernels sciREPL ships with. The full list:

* **R** (webR / WASM) — full R with `install.packages()`, plotting, and SharedVFS * **Python** (Pyodide) — NumPy and SymPy preloaded, `%pip install` for pure-Python packages * **Prolog** (swipl-wasm) — full SWI-Prolog * **ClojureScript** (Scittle v0.6.22) — SCI-based, no build step; also a Prolog code generation target via UnifyWeaver *(experimental)* * **Bash** (brush-wasm) — Unix shell with coreutils, findutils, and grep * **JavaScript** — native browser execution, zero download * **Lua** (Fengari) — lightweight, with notebook cell access via `nb.read()` / `nb.write()`

All kernels share a **SharedVFS** in-memory filesystem, so data flows freely between languages in the same notebook.

**Free & open-source version**

Available as a PWA and an Android app:

* PWA: [https://s243a.github.io/SciREPL/\](https://s243a.github.io/SciREPL/) * Source: [https://github.com/s243a/SciREPL\](https://github.com/s243a/SciREPL)

**Android Pro version**

An Android-only Pro version adds AI-assisted building and bundles more packages locally (with CDN fallback for the rest). AI is integrated via direct API key connections to models, as well as remote connections to coding agents (e.g. Claude Code and Codex).

**Data privacy**

We store no data on an external server. All data is stored locally on your device; packages not bundled with the app are fetched from the CDN when needed. The app is Capacitor-based, so it runs with web browser-level sandboxing. JavaScript, Bash, and TypeR are always offline. Prolog, Python, and ClojureScript are bundled in both free and Pro versions. Lua fetches from CDN on first use. The Pro version also bundles R locally; the free version fetches the R runtime from the CDN (\~50 MB, cached after first use). Note that additional R packages (e.g. tidyverse, ggplot2) are not pre-installed in either version — they can be installed at runtime via `install.packages()`, fetching from the CDN as needed.

**Become a tester**

I need testers before submitting either version to the Play Store. Testers get early access to both the free and Pro versions at no cost. Reply here or DM me if you're interested. In the meantime, you can try the free version in your browser or sideload it via adb.

reddit.com
u/s243a — 1 month ago
▲ 5 r/rstats

[sciREPL v1.0.1] Substantially improved TypeR support — testers wanted for Play Store launch

I'm releasing sciREPL v1.0.1, which brings major improvements to TypeR integration. sciREPL is a Jupyter-style notebook app that lets you mix multiple languages in a single worksheet.

Why we forked TypeR

The standard TypeR type system couldn't express variable-argument (variadic) inputs in its standard-library declarations. We extended it to properly type variadic R functions like cat, paste, sprintf, and c.

What's new in v1.0.1

  • Direct TypeR notebook cells — no longer need to wrap most code in R blocks
  • Named and heterogeneous variadic arguments, including forwarding collected args into another variadic call
  • Named #!source cells for storing reusable TypeR source without executing it or producing output
  • Prolog → TypeR compilation, with generated code placed in a named cell ready for execution
  • Two included typR workbooks: "TypeR Introduction" and "Prolog Generates TypeR" — both accessible via Browse Packages, Bundles & Workbooks in the sciREPL menu, along with workbooks for other supported languages

Other supported languages

TypeR is one of several kernels sciREPL ships with. The full list:

  • R (webR / WASM) — full R with install.packages(), plotting, and SharedVFS
  • Python (Pyodide) — NumPy and SymPy preloaded, %pip install for pure-Python packages
  • Prolog (swipl-wasm) — full SWI-Prolog
  • ClojureScript (Scittle v0.6.22) — SCI-based, no build step; also a Prolog code generation target via UnifyWeaver (experimental)
  • Bash (brush-wasm) — Unix shell with coreutils, findutils, and grep
  • JavaScript — native browser execution, zero download
  • Lua (Fengari) — lightweight, with notebook cell access via nb.read() / nb.write()

All kernels share a SharedVFS in-memory filesystem, so data flows freely between languages in the same notebook.

Free & open-source version

Available as a PWA and an Android app:

Android Pro version

An Android-only Pro version adds AI-assisted building and bundles more packages locally (with CDN fallback for the rest). AI is integrated via direct API key connections to models, as well as remote connections to coding agents (e.g. Claude Code and Codex).

Data privacy

We store no data on an external server. All data is stored locally on your device; packages not bundled with the app are fetched from the CDN when needed. The app is Capacitor-based, so it runs with web browser-level sandboxing. JavaScript, Bash, and TypeR are always offline. Prolog, Python, and ClojureScript are bundled in both free and Pro versions. Lua fetches from CDN on first use. The Pro version also bundles R locally; the free version fetches the R runtime from the CDN (~50 MB, cached after first use). Note that additional R packages (e.g. tidyverse, ggplot2) are not pre-installed in either version — they can be installed at runtime via install.packages(), fetching from the CDN as needed.

Become a tester

I need testers before submitting either version to the Play Store. Testers get early access to both the free and Pro versions at no cost. Reply here or DM me if you're interested. In the meantime, you can try the free version in your browser or sideload it via adb.

reddit.com
u/s243a — 1 month ago

Since this seems to be a complaining form, let me say, "oh no I can't use my gpt subscription for 5 days", because I exhausted my weekly usage! It is April 30th and here is the message,

" You've hit your usage limit. Upgrade to Pro (https://chatgpt.com/explore/pro), visit

https://chatgpt.com/codex/settings/usage to purchase more credits or try again at May 5th, 2026 5:40 AM"

and ironically, my Claude code usage lasted me a whole week. Is it a fair comparison? In this case maybe, sort of. I'm on the 100USD/m plan at anthropic and I have two 20USD/month plans at openAI. I use my anthropic plan twice as much as my openAI account and Claude code is faster than codex/gpt, so for reasons of speed alone we expect claude to burn more tokens. Also claude has longer context length, so if you let your conversations run longer you'll burn tokens faster with anthropic then open AI. Either start new conversations or compact. For best value keep your conversations under 200k tokens, I usually compact around 400k, Anthropic models support a million tokens.

If we look at the API pricing:

"GPT-5.5 charges $5 per 1M input tokens and $30 per 1M output tokens, while Claude Opus 4.7 charges $5 per 1M input tokens and $25 per 1M output tokens, with a  surcharge for prompts exceeding 200K tokens."

Anthropic is actually cheaper if you don't exceed 200k tokens. So if the API prices are similar, I don't expect huge differences in the subscription values. GPT might be say 25% better value but this isn't obvious to me. I find gemini the worse value but you never here people complain about it, because the real battle is between openAI and anthropic, and I wonder how much of the negative sentiment here is bot driven!

reddit.com
u/s243a — 4 months ago