LocalGemmaChat — Help & Support

LocalGemmaChat is a fully on-device AI chat app for iPhone. The language model runs on your own hardware: no account, no sign-in, no cloud service, and nothing about your conversations leaves your device. This page explains how to use the app, what its limits are, and how to fix common problems.

Version documented1.1.5 (build 1)
Last updatedOctober 4, 2026
PlatformiPhone, iOS 26.0 or later
ModelGemma 4 E2B Instruct (text-only, 4-bit) — about 2.5 GiB, downloaded once
Internet neededOnly for the first model download. Fully offline afterwards.
PriceFree — no purchases, no subscriptions, no other charges (since October 4, 2026)
Support contactvegaodm@outlook.com

Getting started

  1. Check your device. You need an iPhone with iOS 26.0 or later and roughly 3 GB of free storage (2.5 GiB for the model, plus working space).
  2. Download the model. On first launch the app asks you to confirm the download and shows its size. Accept, keep the app in the foreground on a reliable connection, and wait. The download uses a background session, so it can continue if you briefly leave the app.
  3. Start chatting. Once the model is loaded you can type a message and send it. From this point on the app works entirely offline — you can put the phone in Airplane Mode and it will still reply.

Model menu (bottom of the screen)

“…” menu (top-right)

Features

Tools calling (recommended for anything involving numbers)

When enabled, the model may call built-in tools instead of answering from memory. This makes arithmetic reliable, because the calculation is performed by the app rather than guessed by the model:

Tools run locally, so enabling them does not affect privacy. Enabling tools makes each reply slightly slower, because the model takes an extra round to request the tool and then read its result.

Thinking

Thinking turns on the model’s built-in reasoning mode, in which it works through a problem before answering. It is off by default because it is noticeably slower and it consumes part of the same limited reply budget described below. Your choice is remembered between launches. Because the setting is applied when the model session is created, switching it reloads the model — expect a few seconds of “Loading model weights…” and a brief full-screen overlay. Your conversation is preserved across the reload.

Two toggles you may want to combine: leave Tools calling on if you ask questions with numbers, and leave Thinking off unless you specifically want the model to reason step by step. Both are independent.

How the conversation and its limits work

LocalGemmaChat runs a small 4-bit model that must fit in your iPhone’s memory, so it deliberately keeps very little context. Understanding this explains most surprising behaviour:

Troubleshooting

SymptomWhat to do
The app closes by itself during a long reply This is iOS terminating the app for exceeding memory or CPU limits while generating a long response. Keep replies shorter, turn Thinking off, and avoid asking for very long texts (for example multi-page documents). Your saved history is preserved, so reopen the app and continue — the conversation is still there.
A reply stops mid-sentence Normal: the reply hit the ~1,024-token limit. Send “continue” to get the rest.
The model seems to forget what we discussed Expected. Only the last two messages are sent each turn. For anything that must be remembered, repeat it in your current message.
The download stalls or fails Confirm you have a working connection and free space, then try again — the download resumes and the app re-checks the files. If it keeps failing, switch HF Source to the other host.
“Loading model weights…” seems to take a long time The app is loading roughly 2.5 GiB of weights into memory. A cold load takes noticeably longer than a reload of an already-downloaded model. Wait for the overlay to disappear; do not force-quit.
You need free space Open the model menu and choose DEL <model>. This removes the weights (about 2.5 GiB) and that model’s history. Everything else stays.
The app reports that the model is not loaded This existed in earlier builds and is fixed in 1.1.2: clearing the chat no longer unloads the model. If you still see it, close and reopen the app.
You want to start over completely Use Clear Chat (Storage screen) to wipe the conversation. To remove everything — model, history, preferences and the local diagnostics log — delete the app from the Home Screen.

Frequently asked questions

Do I need an internet connection?

Only to download the model (about 2.5 GiB) and to check its file sizes beforehand. After that the app is fully usable offline.

Is my conversation sent anywhere?

No. The model runs on your iPhone. There is no account, no analytics, no advertising, no crash reporting and no tracking. The only network requests the app ever makes are the model download and the small metadata request that tells it the download size. See the Privacy Policy for the exact details.

Are there subscriptions, in-app purchases or premium features?

No. LocalGemmaChat is free: there are no in-app purchases, no subscriptions and no other charges, and every feature is available at no cost (as of October 4, 2026). If paid features are ever introduced, the pricing details will be updated on this page and in the Privacy Policy first.

Which AI model does it use?

Gemma 4 E2B Instruct, a text-only 4-bit build published on Hugging Face as over-show/gemma-4-e2b-it-text-only-4bit. It is distributed under Google’s Gemma Terms of Use, so Google’s terms and its Prohibited Use Policy apply to your use of the model’s outputs, in addition to these support notes. This app is not affiliated with or endorsed by Google. Only one model ships with the app, and the app has no interface for adding your own.

Can I send images?

No. The bundled model is text-only, so the app does not offer image input.

Where is my data stored, and how do I delete it?

Everything is inside the app’s private storage on your iPhone: the model in Documents/.localllmclient/, your chat history in Application Support/LLMclient/History/, one preference in system settings storage, and a small local memory diagnostics log in Documents/diagnostics.log. Use Clear Chat (Storage screen) for the conversation, DEL <model> for the model files, or delete the app to remove all of it. Nothing is stored on any server.

Why is the app “LocalGemmaChat” but the bundle identifier says “LLMclient”?

The app was renamed to LocalGemmaChat in version 1.1.2. Its technical bundle identifier was deliberately left unchanged so that existing installations keep their downloaded model, history and preferences.

Is the weather tool giving me real weather?

No. It is a demonstration tool that returns fixed sample data so you can see tool calling at work. It never contacts a weather service.

What changed in 1.1.3?

What changed in 1.1.2?

Reporting a problem

Email vegaodm@outlook.com. To help diagnose the issue quickly, please include:

LocalGemmaChat — ヘルプとサポート

LocalGemmaChat は、完全に端末上で動作する iPhone 向けの AI チャットアプリです。言語モデルはお使いのハードウェア上で動作し、アカウントもログインもクラウドサービスもなく、会話の内容が端末の外に出ることはありません。このページでは、アプリの使い方、制限事項、よくある問題の対処方法を説明します。

対象バージョン1.1.5(build 1)
最終更新2026年10月4日
プラットフォームiPhone、iOS 26.0 以降
モデルGemma 4 E2B Instruct(テキスト専用、4bit)— 約 2.5 GiB、ダウンロードは 1 回のみ
インターネット接続初回のモデルダウンロード時のみ必要です。以後は完全にオフラインで動作します。
料金無料 — 買い切り・サブスクリプション等の課金はありません(2026年10月4日以降)
サポート窓口vegaodm@outlook.com

はじめに

  1. 端末を確認します。 iOS 26.0 以降を実行する iPhone と、およそ 3 GB の空き容量(モデル用の 2.5 GiB と作業領域)が必要です。
  2. モデルをダウンロードします。 初回起動時に、アプリがダウンロードの確認とサイズを表示します。承認し、安定した接続のもとでアプリを前面に保ったまま待ってください。ダウンロードにはバックグラウンドセッションを使うため、アプリを短時間離れても続行されます。
  3. チャットを始めます。 モデルの読み込みが完了したら、メッセージを入力して送信できます。以降は完全にオフラインで動作するため、機内モードにしても返答します。

モデルメニュー(画面下部)

「…」メニュー(右上)

機能

Tools calling(数値を扱う質問におすすめ)

有効にすると、モデルは記憶から答える代わりに組み込みツールを呼び出せるようになります。計算はモデルが推測するのではなくアプリが実行するため、算術が確実になります。

ツールはローカルで動作するため、有効にしてもプライバシーには影響しません。ただし、モデルがツールを要求し、その結果を読み取る分だけ余分な処理が入るため、各返答は少し遅くなります。

Thinking

Thinking は、答える前に問題をじっくり考える、モデル内蔵の推論モードを有効にします。既定ではオフです。体感できるほど遅くなり、後述する限られた返答の予算を一部消費するためです。設定は起動のたびに保持されます。この設定はモデルのセッション作成時に適用されるため、切り替えるとモデルが再読み込みされます。数秒間の「Loading model weights…」と短い全画面表示が入りますが、会話は保持されます。

組み合わせて使いたい 2 つの切り替え: 数値を含む質問をする場合は Tools calling をオンのままにし、モデルに手順を追って考えてほしい場合を除いて Thinking はオフのままにしてください。両者は独立しています。

会話とその制限の仕組み

LocalGemmaChat は iPhone のメモリに収める必要がある小さな 4bit モデルを実行するため、意図的にごくわずかな文脈しか保持しません。これを理解すると、意外に思える挙動のほとんどが説明できます。

トラブルシューティング

症状対処方法
長い返答の途中でアプリが勝手に終了する 長い回答を生成中にメモリまたは CPU の上限を超えたため、iOS がアプリを終了させています。返答を短めにし、Thinking をオフにし、非常に長いテキスト(複数ページの文書など)を求めないようにしてください。保存済みの履歴は残るので、アプリを開き直せば会話はそのまま続けられます。
返答が文の途中で止まる 正常です。返答が約 1,024 トークンの上限に達したためです。「continue」と送ると続きが得られます。
話した内容をモデルが忘れているようだ 想定どおりの動作です。各ターンで送られるのは直近 2 件のメッセージだけです。覚えていてほしい内容は、そのときのメッセージにもう一度書いてください。
ダウンロードが止まる、または失敗する 接続と空き容量を確認してから再試行してください。ダウンロードは再開し、アプリがファイルを再確認します。それでも失敗する場合は、HF Source を別のホストに切り替えてください。
「Loading model weights…」が長く続く 約 2.5 GiB の重みをメモリに読み込んでいます。初回の読み込みは、ダウンロード済みモデルの再読み込みより明らかに時間がかかります。画面が消えるまで待ち、強制終了しないでください。
空き容量を確保したい モデルメニューを開き DEL <モデル> を選びます。重み(約 2.5 GiB)とそのモデルの履歴が削除されます。それ以外は残ります。
モデルが読み込まれていないという表示が出る 以前のビルドにあった問題で、1.1.2 で修正されています(チャットを消去してもモデルがアンロードされなくなりました)。まだ表示される場合は、アプリを終了して開き直してください。
完全に最初からやり直したい 会話を消すには Clear Chat(Storage 画面)を使います。モデル、履歴、設定、ローカルの診断ログまですべて削除するには、ホーム画面からアプリを削除してください。

よくある質問

インターネット接続は必要ですか?

モデルのダウンロード(約 2.5 GiB)と、その前にファイルサイズを確認するときだけ必要です。それ以降は完全にオフラインで利用できます。

会話はどこかに送信されますか?

いいえ。モデルはお使いの iPhone 上で動作します。アカウント、解析、広告、クラッシュレポート、トラッキングはありません。アプリが行うネットワーク通信は、モデルのダウンロードと、ダウンロードサイズを取得する小さなメタデータ要求だけです。詳しくはプライバシーポリシーをご覧ください。

サブスクリプション、アプリ内課金、有料機能はありますか?

いいえ。LocalGemmaChat は無料です。アプリ内課金もサブスクリプションもその他の費用もなく、すべての機能を無償で利用できます(2026年10月4日から)。将来、有料機能を導入する場合は、まずこのページとプライバシーポリシーで内容を更新します。

どの AI モデルを使っていますか?

Gemma 4 E2B Instruct です。Hugging Face で over-show/gemma-4-e2b-it-text-only-4bit として公開されているテキスト専用の 4bit 版です。Google の Gemma 利用規約に基づいて配布されているため、モデルの出力を利用する際には、このサポートページに加えて同規約と禁止使用ポリシーが適用されます。本アプリは Google と提携しておらず、承認も受けていません。同梱されるモデルは 1 つだけで、独自のモデルを追加するインターフェースはありません。

画像を送信できますか?

いいえ。同梱のモデルはテキスト専用のため、画像入力には対応していません。

データはどこに保存され、どうやって削除しますか?

すべて iPhone 上の本アプリ専用領域にあります。モデルは Documents/.localllmclient/、チャット履歴は Application Support/LLMclient/History/、設定はシステムの設定ストレージに 1 件、そして小さなローカルのメモリ診断ログが Documents/diagnostics.log に保存されます。会話は Clear Chat(Storage 画面)、モデルファイルは DEL <モデル>、すべてを削除するにはアプリの削除を使ってください。サーバーには何も保存されていません。

アプリ名は「LocalGemmaChat」なのに、バンドル ID が「LLMclient」なのはなぜですか?

バージョン 1.1.2 でアプリ名を LocalGemmaChat に変更しました。技術的なバンドル ID は、既存のインストール環境でダウンロード済みのモデル、履歴、設定をそのまま引き継げるよう、意図的に変更していません。

天気ツールは実際の天気を返しますか?

いいえ。ツール呼び出しの動作を確認できるよう、固定のサンプルデータを返すデモ用ツールです。気象サービスには一切接続しません。

1.1.3 で何が変わりましたか?

1.1.2 で何が変わりましたか?

問題の報告

vegaodm@outlook.com までメールでお寄せください。迅速に診断するため、次の情報を添えてください。

LocalGemmaChat — Hilfe & Support

LocalGemmaChat ist eine vollständig auf dem Gerät arbeitende KI-Chat-App für das iPhone. Das Sprachmodell läuft auf Ihrer eigenen Hardware: kein Konto, keine Anmeldung, kein Cloud-Dienst, und nichts von Ihren Gesprächen verlässt Ihr Gerät. Diese Seite erklärt, wie Sie die App bedienen, welche Grenzen sie hat und wie Sie häufige Probleme beheben.

Dokumentierte Version1.1.5 (Build 1)
Zuletzt aktualisiert4. Oktober 2026
PlattformiPhone, iOS 26.0 oder neuer
ModellGemma 4 E2B Instruct (nur Text, 4 Bit) — etwa 2,5 GiB, einmaliger Download
Internet erforderlichNur für den ersten Modell-Download. Danach vollständig offline.
PreisKostenlos — keine Käufe, keine Abonnements, keine sonstigen Kosten (seit 4. Oktober 2026)
Support-Kontaktvegaodm@outlook.com

Erste Schritte

  1. Prüfen Sie Ihr Gerät. Sie benötigen ein iPhone mit iOS 26.0 oder neuer und rund 3 GB freien Speicher (2,5 GiB für das Modell plus Arbeitsplatz).
  2. Laden Sie das Modell herunter. Beim ersten Start bittet die App um Bestätigung des Downloads und zeigt dessen Größe an. Bestätigen Sie, lassen Sie die App bei einer zuverlässigen Verbindung im Vordergrund und warten Sie. Der Download nutzt eine Hintergrundsitzung und kann daher fortgesetzt werden, wenn Sie die App kurz verlassen.
  3. Beginnen Sie zu chatten. Sobald das Modell geladen ist, können Sie eine Nachricht eingeben und senden. Ab diesem Zeitpunkt arbeitet die App völlig offline — Sie können den Flugmodus aktivieren und sie antwortet weiterhin.

Modellmenü (unten am Bildschirm)

Menü „…“ (oben rechts)

Funktionen

Tools calling (empfohlen bei allem mit Zahlen)

Wenn aktiviert, kann das Modell integrierte Werkzeuge aufrufen, statt aus dem Gedächtnis zu antworten. Das macht Arithmetik zuverlässig, weil die Berechnung von der App ausgeführt und nicht vom Modell geraten wird:

Die Werkzeuge laufen lokal, die Aktivierung beeinträchtigt also den Datenschutz nicht. Sie machen jede Antwort etwas langsamer, weil das Modell eine zusätzliche Runde benötigt, um das Werkzeug anzufordern und das Ergebnis zu lesen.

Thinking

Thinking aktiviert den eingebauten Denkmodus des Modells, in dem es ein Problem durcharbeitet, bevor es antwortet. Er ist standardmäßig aus, weil er spürbar langsamer ist und einen Teil desselben begrenzten Antwortbudgets verbraucht, das weiter unten beschrieben wird. Ihre Wahl bleibt über Neustarts hinweg erhalten. Da die Einstellung beim Erstellen der Modellsitzung angewendet wird, lädt ein Umschalten das Modell neu — rechnen Sie mit einigen Sekunden „Loading model weights…“ und einer kurzen Vollbildanzeige. Ihr Gespräch bleibt über das Neuladen hinweg erhalten.

Zwei Schalter, die Sie kombinieren können: Lassen Sie Tools calling aktiv, wenn Sie Fragen mit Zahlen stellen, und lassen Sie Thinking aus, wenn Sie nicht ausdrücklich möchten, dass das Modell Schritt für Schritt denkt. Beide sind unabhängig.

Wie das Gespräch und seine Grenzen funktionieren

LocalGemmaChat führt ein kleines 4-Bit-Modell aus, das in den Speicher Ihres iPhones passen muss, und hält daher bewusst sehr wenig Kontext vor. Wer das versteht, kann die meisten überraschenden Verhaltensweisen erklären:

Fehlerbehebung

SymptomWas zu tun ist
Die App schließt sich während einer langen Antwort von selbst Hier beendet iOS die App, weil beim Erzeugen einer langen Antwort Speicher- oder CPU-Grenzen überschritten werden. Halten Sie Antworten kürzer, schalten Sie Thinking aus und verlangen Sie keine sehr langen Texte (etwa mehrseitige Dokumente). Ihr gespeicherter Verlauf bleibt erhalten: Öffnen Sie die App erneut und machen Sie weiter — das Gespräch ist noch da.
Eine Antwort bricht mitten im Satz ab Normal: Die Antwort hat die Grenze von ~1.024 Token erreicht. Senden Sie „continue“, um den Rest zu erhalten.
Das Modell scheint zu vergessen, was besprochen wurde Erwartet. Pro Runde werden nur die letzten zwei Nachrichten gesendet. Wiederholen Sie alles, was im Gedächtnis bleiben muss, in Ihrer aktuellen Nachricht.
Der Download hängt oder schlägt fehl Prüfen Sie Verbindung und freien Speicher und versuchen Sie es erneut — der Download wird fortgesetzt und die App prüft die Dateien erneut. Wenn es weiterhin fehlschlägt, stellen Sie HF Source auf den anderen Host um.
„Loading model weights…“ dauert scheinbar lange Die App lädt rund 2,5 GiB Gewichte in den Speicher. Ein Kaltstart dauert deutlich länger als das Neuladen eines bereits heruntergeladenen Modells. Warten Sie, bis die Anzeige verschwindet; beenden Sie die App nicht gewaltsam.
Sie benötigen freien Speicher Öffnen Sie das Modellmenü und wählen Sie DEL <Modell>. Damit werden die Gewichte (etwa 2,5 GiB) und der Verlauf dieses Modells entfernt. Alles andere bleibt.
Die App meldet, das Modell sei nicht geladen Das gab es in früheren Builds und ist in 1.1.2 behoben: Das Löschen des Chats entlädt das Modell nicht mehr. Sollte es weiterhin auftreten, schließen Sie die App und öffnen Sie sie erneut.
Sie möchten komplett neu beginnen Mit Clear Chat (Storage-Bildschirm) löschen Sie das Gespräch. Um alles zu entfernen — Modell, Verlauf, Einstellungen und das lokale Diagnoseprotokoll — löschen Sie die App vom Home-Bildschirm.

Häufige Fragen

Brauche ich eine Internetverbindung?

Nur zum Herunterladen des Modells (etwa 2,5 GiB) und um vorab dessen Dateigrößen zu prüfen. Danach ist die App vollständig offline nutzbar.

Wird mein Gespräch irgendwohin gesendet?

Nein. Das Modell läuft auf Ihrem iPhone. Es gibt kein Konto, keine Analyse, keine Werbung, keine Absturzberichte und kein Tracking. Die einzigen Netzwerkanfragen der App sind der Modell-Download und die kleine Metadatenanfrage, die die Downloadgröße ermittelt. Die genauen Details finden Sie in der Datenschutzerklärung.

Gibt es Abonnements, In-App-Käufe oder Premium-Funktionen?

Nein. LocalGemmaChat ist kostenlos: Es gibt keine In-App-Käufe, keine Abonnements und keine sonstigen Kosten, und alle Funktionen sind kostenlos nutzbar (gültig ab dem 4. Oktober 2026). Sollten jemals kostenpflichtige Funktionen eingeführt werden, werden die Preisangaben zuerst auf dieser Seite und in der Datenschutzerklärung aktualisiert.

Welches KI-Modell wird verwendet?

Gemma 4 E2B Instruct, ein reiner 4-Bit-Textaufbau, veröffentlicht auf Hugging Face als over-show/gemma-4-e2b-it-text-only-4bit. Es wird unter Googles Gemma-Nutzungsbedingungen verbreitet; daher gelten Googles Bedingungen und seine Richtlinie zur verbotenen Nutzung zusätzlich zu diesen Support-Hinweisen für Ihre Verwendung der Modellausgaben. Diese App ist mit Google weder verbunden noch von Google unterstützt. Es ist nur ein Modell enthalten, und die App bietet keine Schnittstelle zum Hinzufügen eigener Modelle.

Kann ich Bilder senden?

Nein. Das mitgelieferte Modell ist reiner Text, daher bietet die App keine Bildeingabe an.

Wo werden meine Daten gespeichert, und wie lösche ich sie?

Alles liegt im privaten Speicher der App auf Ihrem iPhone: das Modell in Documents/.localllmclient/, Ihr Chatverlauf in Application Support/LLMclient/History/, eine Einstellung im Systemspeicher und ein kleines lokales Speicherdiagnoseprotokoll in Documents/diagnostics.log. Nutzen Sie Clear Chat (Storage-Bildschirm) für das Gespräch, DEL <Modell> für die Modelldateien, oder löschen Sie die App, um alles zu entfernen. Auf keinem Server wird etwas gespeichert.

Warum heißt die App „LocalGemmaChat“, der Bundle-Identifier aber „LLMclient“?

Die App wurde in Version 1.1.2 in LocalGemmaChat umbenannt. Ihr technischer Bundle-Identifier wurde bewusst unverändert gelassen, damit bestehende Installationen ihr heruntergeladenes Modell, ihren Verlauf und ihre Einstellungen behalten.

Liefert das Wetter-Werkzeug echtes Wetter?

Nein. Es ist ein Demonstrationswerkzeug, das feste Beispieldaten zurückgibt, damit Sie den Werkzeugaufruf in Aktion sehen. Es kontaktiert nie einen Wetterdienst.

Was hat sich in 1.1.3 geändert?

Was hat sich in 1.1.2 geändert?

Ein Problem melden

Schreiben Sie an vegaodm@outlook.com. Um das Problem schnell einzugrenzen, geben Sie bitte an:

LocalGemmaChat — Aide et assistance

LocalGemmaChat est une application de conversation par IA entièrement embarquée pour iPhone. Le modèle de langage s’exécute sur votre propre matériel : aucun compte, aucune connexion, aucun service cloud, et rien de vos conversations ne quitte votre appareil. Cette page explique comment utiliser l’application, quelles sont ses limites et comment résoudre les problèmes courants.

Version documentée1.1.5 (build 1)
Dernière mise à jour4 octobre 2026
PlateformeiPhone, iOS 26.0 ou version ultérieure
ModèleGemma 4 E2B Instruct (texte uniquement, 4 bits) — environ 2,5 GiB, téléchargé une seule fois
Internet nécessaireUniquement pour le premier téléchargement du modèle. Entièrement hors ligne ensuite.
PrixGratuit — aucun achat, aucun abonnement, aucun autre frais (depuis le 4 octobre 2026)
Contact du supportvegaodm@outlook.com

Pour commencer

  1. Vérifiez votre appareil. Il vous faut un iPhone sous iOS 26.0 ou version ultérieure et environ 3 Go d’espace libre (2,5 GiB pour le modèle, plus de l’espace de travail).
  2. Téléchargez le modèle. Au premier lancement, l’application vous demande de confirmer le téléchargement et en affiche la taille. Acceptez, gardez l’application au premier plan sur une connexion fiable et attendez. Le téléchargement utilise une session d’arrière-plan : il peut donc se poursuivre si vous quittez brièvement l’application.
  3. Commencez à discuter. Une fois le modèle chargé, vous pouvez saisir un message et l’envoyer. À partir de là, l’application fonctionne entièrement hors ligne — vous pouvez activer le mode Avion, elle répondra toujours.

Menu du modèle (en bas de l’écran)

Menu « … » (en haut à droite)

Fonctionnalités

Tools calling (recommandé dès qu’il y a des chiffres)

Lorsque cette option est active, le modèle peut appeler des outils intégrés au lieu de répondre de mémoire. L’arithmétique devient fiable, car le calcul est effectué par l’application plutôt que deviné par le modèle :

Les outils s’exécutent localement : les activer ne change rien à la confidentialité. Les activer ralentit légèrement chaque réponse, car le modèle a besoin d’un tour supplémentaire pour demander l’outil puis lire son résultat.

Thinking

Thinking active le mode de raisonnement intégré du modèle, dans lequel il travaille un problème avant de répondre. Il est désactivé par défaut, car il est nettement plus lent et consomme une partie du budget de réponse limité décrit ci-dessous. Votre choix est mémorisé entre les lancements. Comme le réglage est appliqué à la création de la session du modèle, le basculer recharge le modèle — comptez quelques secondes de « Loading model weights… » et un bref écran plein format. Votre conversation est conservée pendant le rechargement.

Deux options à combiner : laissez Tools calling activé si vous posez des questions avec des chiffres, et laissez Thinking désactivé, sauf si vous voulez réellement que le modèle raisonne étape par étape. Les deux sont indépendantes.

Fonctionnement de la conversation et de ses limites

LocalGemmaChat exécute un petit modèle 4 bits qui doit tenir dans la mémoire de votre iPhone ; il ne conserve donc volontairement que très peu de contexte. Le comprendre explique la plupart des comportements surprenants :

Résolution des problèmes

SymptômeQue faire
L’application se ferme seule pendant une longue réponse iOS ferme l’application parce qu’elle dépasse les limites de mémoire ou de CPU pendant la génération d’une longue réponse. Faites des réponses plus courtes, désactivez Thinking et évitez de demander des textes très longs (par exemple des documents de plusieurs pages). Votre historique enregistré est conservé : rouvrez l’application et poursuivez, la conversation est toujours là.
Une réponse s’arrête au milieu d’une phrase Normal : la réponse a atteint la limite d’environ 1 024 tokens. Envoyez « continue » pour obtenir la suite.
Le modèle semble oublier ce dont nous avons parlé C’est attendu. Seuls les deux derniers messages sont envoyés à chaque tour. Pour tout ce qui doit être retenu, répétez-le dans votre message actuel.
Le téléchargement se bloque ou échoue Vérifiez votre connexion et l’espace libre, puis réessayez — le téléchargement reprend et l’application revérifie les fichiers. S’il échoue encore, changez HF Source pour l’autre hôte.
« Loading model weights… » semble très long L’application charge environ 2,5 GiB de poids en mémoire. Un chargement à froid est nettement plus long qu’un rechargement d’un modèle déjà téléchargé. Attendez la disparition de l’écran ; ne forcez pas la fermeture.
Vous avez besoin d’espace libre Ouvrez le menu du modèle et choisissez DEL <modèle>. Cela supprime les poids (environ 2,5 GiB) et l’historique de ce modèle. Tout le reste est conservé.
L’application indique que le modèle n’est pas chargé Cela existait dans les versions antérieures et est corrigé en 1.1.2 : effacer la conversation ne décharge plus le modèle. Si cela se produit encore, fermez puis rouvrez l’application.
Vous voulez repartir de zéro Utilisez Clear Chat (écran Storage) pour effacer la conversation. Pour tout supprimer — modèle, historique, préférences et journal de diagnostic local — supprimez l’application depuis l’écran d’accueil.

Questions fréquentes

Ai-je besoin d’une connexion Internet ?

Uniquement pour télécharger le modèle (environ 2,5 GiB) et vérifier sa taille au préalable. Ensuite, l’application est pleinement utilisable hors ligne.

Ma conversation est-elle envoyée quelque part ?

Non. Le modèle s’exécute sur votre iPhone. Il n’y a ni compte, ni analyse, ni publicité, ni rapport de plantage, ni suivi. Les seules requêtes réseau de l’application sont le téléchargement du modèle et la petite requête de métadonnées qui en donne la taille. Voir la Politique de confidentialité pour les détails exacts.

Y a-t-il des abonnements, des achats intégrés ou des fonctions premium ?

Non. LocalGemmaChat est gratuit : il n’y a ni achats intégrés, ni abonnements, ni aucun autre frais, et toutes les fonctions sont accessibles gratuitement (à compter du 4 octobre 2026). Si des fonctions payantes devaient être introduites un jour, les informations tarifaires seront mises à jour d’abord sur cette page et dans la Politique de confidentialité.

Quel modèle d’IA est utilisé ?

Gemma 4 E2B Instruct, une version texte uniquement en 4 bits publiée sur Hugging Face sous over-show/gemma-4-e2b-it-text-only-4bit. Il est distribué sous les Conditions d’utilisation de Gemma de Google : ces conditions et la politique d’utilisation interdite de Google s’appliquent donc à votre usage des sorties du modèle, en plus de ces notes d’assistance. Cette application n’est ni affiliée à Google ni approuvée par Google. Un seul modèle est fourni, et l’application n’offre aucune interface pour en ajouter.

Puis-je envoyer des images ?

Non. Le modèle fourni est uniquement textuel : l’application ne propose donc pas d’entrée d’images.

Où sont stockées mes données et comment les supprimer ?

Tout se trouve dans l’espace de stockage privé de l’application sur votre iPhone : le modèle dans Documents/.localllmclient/, votre historique de conversation dans Application Support/LLMclient/History/, une préférence dans le stockage système et un petit journal local de diagnostic mémoire dans Documents/diagnostics.log. Utilisez Clear Chat (écran Storage) pour la conversation, DEL <modèle> pour les fichiers du modèle, ou supprimez l’application pour tout effacer. Rien n’est stocké sur un serveur.

Pourquoi l’application s’appelle-t-elle « LocalGemmaChat » alors que l’identifiant de bundle indique « LLMclient » ?

L’application a été renommée LocalGemmaChat dans la version 1.1.2. Son identifiant de bundle technique a été délibérément laissé inchangé afin que les installations existantes conservent leur modèle téléchargé, leur historique et leurs préférences.

L’outil météo donne-t-il la météo réelle ?

Non. C’est un outil de démonstration qui renvoie des données d’exemple fixes pour vous montrer l’appel d’outils en action. Il ne contacte jamais de service météo.

Qu’est-ce qui a changé en 1.1.3 ?

Qu’est-ce qui a changé en 1.1.2 ?

Signaler un problème

Écrivez à vegaodm@outlook.com. Pour nous aider à diagnostiquer rapidement le problème, veuillez inclure :

LocalGemmaChat — Ayuda y soporte

LocalGemmaChat es una aplicación de chat con IA totalmente local para iPhone. El modelo de lenguaje se ejecuta en su propio hardware: sin cuenta, sin inicio de sesión, sin servicio en la nube, y nada de sus conversaciones sale de su dispositivo. Esta página explica cómo usar la aplicación, cuáles son sus límites y cómo resolver los problemas habituales.

Versión documentada1.1.5 (compilación 1)
Última actualización4 de octubre de 2026
PlataformaiPhone, iOS 26.0 o posterior
ModeloGemma 4 E2B Instruct (solo texto, 4 bits) — unos 2,5 GiB, se descarga una sola vez
Internet necesarioSolo para la primera descarga del modelo. Totalmente sin conexión después.
PrecioGratis — sin compras, sin suscripciones y sin ningún otro cargo (desde el 4 de octubre de 2026)
Contacto de soportevegaodm@outlook.com

Primeros pasos

  1. Compruebe su dispositivo. Necesita un iPhone con iOS 26.0 o posterior y unos 3 GB de espacio libre (2,5 GiB para el modelo, más espacio de trabajo).
  2. Descargue el modelo. En el primer inicio, la aplicación le pide que confirme la descarga y muestra su tamaño. Acepte, mantenga la aplicación en primer plano con una conexión fiable y espere. La descarga usa una sesión en segundo plano, por lo que puede continuar si sale brevemente de la aplicación.
  3. Empiece a conversar. Una vez cargado el modelo, puede escribir un mensaje y enviarlo. A partir de ese momento la aplicación funciona totalmente sin conexión: puede activar el modo Avión y seguirá respondiendo.

Menú del modelo (parte inferior de la pantalla)

Menú «…» (arriba a la derecha)

Funciones

Tools calling (recomendado para todo lo que implique números)

Cuando está activado, el modelo puede llamar a herramientas integradas en lugar de responder de memoria. Esto hace fiable la aritmética, porque el cálculo lo realiza la aplicación y no lo adivina el modelo:

Las herramientas se ejecutan localmente, así que activarlas no afecta a la privacidad. Activarlas hace que cada respuesta sea algo más lenta, porque el modelo necesita una ronda adicional para solicitar la herramienta y leer su resultado.

Thinking

Thinking activa el modo de razonamiento integrado del modelo, en el que analiza un problema antes de responder. Está desactivado de forma predeterminada porque es notablemente más lento y consume parte del mismo presupuesto limitado de respuesta que se describe más abajo. Su elección se recuerda entre inicios. Como el ajuste se aplica al crear la sesión del modelo, cambiarlo recarga el modelo: cuente con unos segundos de «Loading model weights…» y una breve pantalla completa. Su conversación se conserva durante la recarga.

Dos opciones que quizá quiera combinar: deje Tools calling activado si hace preguntas con números, y deje Thinking desactivado salvo que quiera específicamente que el modelo razone paso a paso. Ambas son independientes.

Cómo funcionan la conversación y sus límites

LocalGemmaChat ejecuta un pequeño modelo de 4 bits que debe caber en la memoria de su iPhone, por lo que conserva deliberadamente muy poco contexto. Entender esto explica la mayoría de los comportamientos sorprendentes:

Solución de problemas

SíntomaQué hacer
La aplicación se cierra sola durante una respuesta larga iOS está cerrando la aplicación por superar los límites de memoria o CPU al generar una respuesta larga. Haga respuestas más cortas, desactive Thinking y evite pedir textos muy largos (por ejemplo, documentos de varias páginas). Su historial guardado se conserva, así que vuelva a abrir la aplicación y continúe: la conversación sigue ahí.
Una respuesta se corta a mitad de frase Es normal: la respuesta alcanzó el límite de ~1.024 tokens. Envíe «continue» para obtener el resto.
El modelo parece olvidar lo que hemos hablado Es lo esperado. En cada turno solo se envían los dos últimos mensajes. Todo lo que deba recordarse, repítalo en su mensaje actual.
La descarga se atasca o falla Compruebe que tiene conexión y espacio libre, y vuelva a intentarlo: la descarga se reanuda y la aplicación vuelve a comprobar los archivos. Si sigue fallando, cambie HF Source al otro host.
«Loading model weights…» parece tardar mucho La aplicación está cargando unos 2,5 GiB de pesos en memoria. Una carga en frío tarda bastante más que una recarga de un modelo ya descargado. Espere a que desaparezca la pantalla; no fuerce el cierre.
Necesita espacio libre Abra el menú del modelo y elija DEL <modelo>. Esto elimina los pesos (unos 2,5 GiB) y el historial de ese modelo. Todo lo demás se conserva.
La aplicación indica que el modelo no está cargado Esto ocurría en versiones anteriores y está corregido en la 1.1.2: borrar la conversación ya no descarga el modelo. Si aún lo ve, cierre y vuelva a abrir la aplicación.
Quiere empezar de cero por completo Usa Clear Chat (pantalla Storage) para borrar la conversación. Para eliminar todo — modelo, historial, preferencias y el registro de diagnóstico local — elimine la aplicación desde la pantalla de inicio.

Preguntas frecuentes

¿Necesito conexión a Internet?

Solo para descargar el modelo (unos 2,5 GiB) y para comprobar antes sus tamaños de archivo. Después, la aplicación es totalmente utilizable sin conexión.

¿Se envía mi conversación a algún sitio?

No. El modelo se ejecuta en su iPhone. No hay cuenta, ni analítica, ni publicidad, ni informes de fallos, ni seguimiento. Las únicas solicitudes de red que hace la aplicación son la descarga del modelo y la pequeña solicitud de metadatos que indica el tamaño de la descarga. Consulte la Política de privacidad para conocer los detalles exactos.

¿Hay suscripciones, compras integradas o funciones premium?

No. LocalGemmaChat es gratis: no hay compras integradas, ni suscripciones, ni ningún otro cargo, y todas las funciones están disponibles sin coste (a partir del 4 de octubre de 2026). Si en el futuro se introducen funciones de pago, los detalles de precio se actualizarán primero en esta página y en la Política de privacidad.

¿Qué modelo de IA utiliza?

Gemma 4 E2B Instruct, una versión solo de texto en 4 bits publicada en Hugging Face como over-show/gemma-4-e2b-it-text-only-4bit. Se distribuye bajo las Condiciones de uso de Gemma de Google, por lo que dichas condiciones y su Política de usos prohibidos se aplican a su uso de las salidas del modelo, además de estas notas de soporte. Esta aplicación no está afiliada a Google ni cuenta con su respaldo. Solo se incluye un modelo y la aplicación no ofrece ninguna interfaz para añadir el suyo.

¿Puedo enviar imágenes?

No. El modelo incluido es solo de texto, así que la aplicación no ofrece entrada de imágenes.

¿Dónde se guardan mis datos y cómo los elimino?

Todo está en el almacenamiento privado de la aplicación en su iPhone: el modelo en Documents/.localllmclient/, su historial de conversaciones en Application Support/LLMclient/History/, una preferencia en el almacenamiento del sistema y un pequeño registro local de diagnóstico de memoria en Documents/diagnostics.log. Usa Clear Chat (pantalla Storage) para la conversación, DEL <modelo> para los archivos del modelo, o elimine la aplicación para quitarlo todo. No se guarda nada en ningún servidor.

¿Por qué la aplicación se llama «LocalGemmaChat» pero el identificador de paquete dice «LLMclient»?

La aplicación pasó a llamarse LocalGemmaChat en la versión 1.1.2. Su identificador de paquete técnico se dejó deliberadamente sin cambios para que las instalaciones existentes conserven el modelo descargado, el historial y las preferencias.

¿La herramienta del tiempo me da el tiempo real?

No. Es una herramienta de demostración que devuelve datos de ejemplo fijos para que vea la llamada a herramientas en funcionamiento. Nunca contacta con un servicio meteorológico.

¿Qué ha cambiado en la 1.1.3?

¿Qué ha cambiado en la 1.1.2?

Informar de un problema

Escriba a vegaodm@outlook.com. Para ayudarnos a diagnosticar el problema con rapidez, incluya:

LocalGemmaChat — Assistenza e supporto

LocalGemmaChat è un’app di chat con IA che funziona interamente sul dispositivo, per iPhone. Il modello linguistico viene eseguito sul tuo hardware: nessun account, nessun accesso, nessun servizio cloud e nulla delle tue conversazioni lascia il dispositivo. Questa pagina spiega come usare l’app, quali sono i suoi limiti e come risolvere i problemi più comuni.

Versione documentata1.1.5 (build 1)
Ultimo aggiornamento4 ottobre 2026
PiattaformaiPhone, iOS 26.0 o successivo
ModelloGemma 4 E2B Instruct (solo testo, 4 bit) — circa 2,5 GiB, scaricato una sola volta
Internet necessarioSolo per il primo download del modello. Poi completamente offline.
PrezzoGratuito — nessun acquisto, nessun abbonamento, nessun altro costo (dal 4 ottobre 2026)
Contatto per il supportovegaodm@outlook.com

Per iniziare

  1. Controlla il dispositivo. Serve un iPhone con iOS 26.0 o successivo e circa 3 GB di spazio libero (2,5 GiB per il modello, più lo spazio di lavoro).
  2. Scarica il modello. Al primo avvio l’app chiede di confermare il download e ne mostra la dimensione. Accetta, mantieni l’app in primo piano con una connessione affidabile e attendi. Il download usa una sessione in background, quindi può proseguire se esci brevemente dall’app.
  3. Inizia a chattare. Una volta caricato il modello puoi digitare un messaggio e inviarlo. Da quel momento l’app funziona completamente offline: puoi attivare la modalità Aereo e continuerà a rispondere.

Menu del modello (in basso sullo schermo)

Menu «…» (in alto a destra)

Funzionalità

Tools calling (consigliato per qualsiasi cosa con numeri)

Quando è attivo, il modello può chiamare strumenti integrati invece di rispondere a memoria. Questo rende affidabile l’aritmetica, perché il calcolo lo esegue l’app invece di essere indovinato dal modello:

Gli strumenti vengono eseguiti localmente, quindi attivarli non incide sulla privacy. Attivarli rende ogni risposta un po’ più lenta, perché il modello impiega un giro in più per richiedere lo strumento e leggerne il risultato.

Thinking

Thinking attiva la modalità di ragionamento integrata del modello, in cui elabora un problema prima di rispondere. È disattivata per impostazione predefinita perché è sensibilmente più lenta e consuma parte dello stesso budget di risposta limitato descritto di seguito. La tua scelta viene ricordata tra un avvio e l’altro. Poiché l’impostazione viene applicata al momento della creazione della sessione del modello, cambiarla ricarica il modello: aspettati qualche secondo di «Loading model weights…» e una breve schermata a tutto schermo. La conversazione viene conservata durante la ricarica.

Due opzioni che potresti voler combinare: lascia attivo Tools calling se fai domande con numeri, e lascia disattivato Thinking a meno che tu non voglia specificamente che il modello ragioni passo per passo. Le due opzioni sono indipendenti.

Come funzionano la conversazione e i suoi limiti

LocalGemmaChat esegue un piccolo modello a 4 bit che deve stare nella memoria del tuo iPhone, quindi mantiene deliberatamente pochissimo contesto. Capire questo spiega la maggior parte dei comportamenti inattesi:

Risoluzione dei problemi

SintomoCosa fare
L’app si chiude da sola durante una risposta lunga iOS sta chiudendo l’app perché ha superato i limiti di memoria o CPU mentre generava una risposta lunga. Mantieni le risposte più brevi, disattiva Thinking ed evita di chiedere testi molto lunghi (per esempio documenti di più pagine). La cronologia salvata viene conservata: riapri l’app e continua, la conversazione è ancora lì.
Una risposta si interrompe a metà frase Normale: la risposta ha raggiunto il limite di ~1.024 token. Invia «continue» per ottenere il resto.
Il modello sembra dimenticare ciò di cui abbiamo parlato È previsto. A ogni turno vengono inviati solo gli ultimi due messaggi. Per tutto ciò che deve essere ricordato, ripetilo nel messaggio corrente.
Il download si blocca o non riesce Verifica di avere una connessione funzionante e spazio libero, poi riprova: il download riprende e l’app ricontrolla i file. Se continua a fallire, imposta HF Source sull’altro host.
«Loading model weights…» sembra durare molto L’app sta caricando in memoria circa 2,5 GiB di pesi. Un caricamento a freddo richiede molto più tempo del ricaricamento di un modello già scaricato. Attendi che la schermata scompaia; non forzare la chiusura.
Ti serve spazio libero Apri il menu del modello e scegli DEL <modello>. Rimuove i pesi (circa 2,5 GiB) e la cronologia di quel modello. Tutto il resto resta.
L’app segnala che il modello non è caricato Esisteva nelle build precedenti ed è stato corretto nella 1.1.2: cancellare la chat non scarica più il modello. Se lo vedi ancora, chiudi e riapri l’app.
Vuoi ricominciare completamente da zero Usa Clear Chat (schermata Storage) per azzerare la conversazione. Per rimuovere tutto — modello, cronologia, preferenze e il registro di diagnostica locale — elimina l’app dalla schermata Home.

Domande frequenti

Serve una connessione a Internet?

Solo per scaricare il modello (circa 2,5 GiB) e per verificarne prima le dimensioni dei file. Dopo, l’app è pienamente utilizzabile offline.

La mia conversazione viene inviata da qualche parte?

No. Il modello viene eseguito sul tuo iPhone. Non ci sono account, analisi, pubblicità, segnalazione di crash né tracciamento. Le uniche richieste di rete che l’app effettua sono il download del modello e la piccola richiesta di metadati che ne indica la dimensione. Consulta l’Informativa sulla privacy per i dettagli esatti.

Ci sono abbonamenti, acquisti in-app o funzioni premium?

No. LocalGemmaChat è gratuito: non ci sono acquisti in-app, abbonamenti né altri costi, e tutte le funzioni sono disponibili gratuitamente (a partire dal 4 ottobre 2026). Se in futuro verranno introdotte funzioni a pagamento, i dettagli sui prezzi saranno aggiornati prima su questa pagina e nell’Informativa sulla privacy.

Quale modello di IA utilizza?

Gemma 4 E2B Instruct, una build solo testo a 4 bit pubblicata su Hugging Face come over-show/gemma-4-e2b-it-text-only-4bit. È distribuito secondo i Termini di utilizzo di Gemma di Google, quindi quei termini e la relativa Politica sui divieti d’uso si applicano al tuo utilizzo degli output del modello, oltre a queste note di supporto. Questa app non è affiliata a Google né approvata da Google. È incluso un solo modello e l’app non offre alcuna interfaccia per aggiungerne altri.

Posso inviare immagini?

No. Il modello incluso è solo testo, quindi l’app non offre input per le immagini.

Dove sono conservati i miei dati e come li elimino?

Tutto si trova nello spazio di archiviazione privato dell’app sul tuo iPhone: il modello in Documents/.localllmclient/, la cronologia delle chat in Application Support/LLMclient/History/, una preferenza nella memoria di sistema e un piccolo registro locale di diagnostica della memoria in Documents/diagnostics.log. Usa Clear Chat (schermata Storage) per la conversazione, DEL <modello> per i file del modello, oppure elimina l’app per rimuovere tutto. Nulla viene conservato su alcun server.

Perché l’app si chiama «LocalGemmaChat» ma l’identificatore del bundle dice «LLMclient»?

L’app è stata rinominata LocalGemmaChat nella versione 1.1.2. Il suo identificatore di bundle tecnico è stato lasciato intenzionalmente invariato, così le installazioni esistenti mantengono il modello scaricato, la cronologia e le preferenze.

Lo strumento meteo fornisce il meteo reale?

No. È uno strumento dimostrativo che restituisce dati di esempio fissi, così puoi vedere la chiamata agli strumenti in azione. Non contatta mai alcun servizio meteo.

Cosa è cambiato nella 1.1.3?

Cosa è cambiato nella 1.1.2?

Segnalare un problema

Scrivi a vegaodm@outlook.com. Per aiutarci a diagnosticare rapidamente il problema, includi: