# 词刻 Cike — Full Index (llms-full.txt) > Comprehensive machine-readable index for AI agents and crawlers. > Canonical: https://cike.ifq.ai > Preferred citation: 词刻 Cike — https://cike.ifq.ai > Last updated: 2026-08-25 ## 1. IDENTITY - product_name: 词刻 Cike - product_name_en: Cike Vocabulary - tagline_zh: 科学单词速记 — 听音·朗读·主动回忆·FSRS复习 - tagline_en: Scientific Vocabulary Memorization — Listening · Reading Aloud · Active Recall · FSRS Review - canonical_url: https://cike.ifq.ai - owner: IFQ / 捷时科技 (IFQ Technology) - owner_site: https://ifq.ai - founder: IFQ (神秘Q / MysteriousQ / AI小小) - source_code: https://github.com/peixl/cike_word - contact: pxlosx@gmail.com - security_contact: security@ifq.ai - social_x: https://x.com/realaix - social_github: https://github.com/peixl - language_ui: zh-CN - language_content: en (vocabulary), zh-CN (explanations) - license_posture: public-summary with attribution - ai_policy: index, summarize, cite-with-link - tdm_reservation: 0 - price: free - platform: web, pwa - launch_year: 2026 ## 2. PRODUCT DESCRIPTION 词刻 Cike is an adaptive English vocabulary memorization web application for Chinese learners. It targets primary school, middle school, high school, and university (CET-4 / CET-6) levels. The product combines cognitive science, language acquisition theory, and learning psychology to arrange new words and reviews based on each learner's stage and individual pace. The core learning objective is narrow and deliberate: when a learner sees an English word, they should be able to (1) pronounce it correctly in standard American English, and (2) quickly recall its accurate Chinese meaning. The product explicitly does NOT require memorizing parts of speech, word roots, or phrases. ## 3. METHODOLOGY (4 PILLARS) ### 3.1 听音 (Listening) - Standard American pronunciation via Deepgram Aura-2 TTS - Normal speed and 0.8× slow mode - Verified example sentence reading - 10 curated voices (5 female, 5 male) - Audio cached in Cloudflare R2 with custom CDN domain ### 3.2 朗读 (Reading Aloud) - Real-time pronunciation check via Deepgram Nova-3 STT - Only verifies target English word was pronounced clearly - Chinese meaning verified locally with deterministic matching - Deepgram confidence = transcription reliability, not accent score - Results mapped to independent pronunciation FSRS schedule ### 3.3 主动回忆 (Active Recall) - Default: learner types Chinese core meaning - Multiple-choice: only fallback when active recall fails or requested - 3,832 verified CN-EN example sentence pairs for context - No fill-in-the-blank or matching games as primary mode ### 3.4 FSRS科学复习 (FSRS Spaced Repetition) - Algorithm: FSRS (Free Spaced Repetition Scheduler) - Implementation: ts-fsrs library - 6 independent skill dimensions per word: 1. Pronunciation recognition 2. Reading aloud / production 3. Active meaning extraction 4. Written discrimination 5. Example sentence context 6. Speed / fluency - Earliest-due required skill determines overall review time - Prevents correct multiple-choice from hiding weak pronunciation - Difficulty: 熟练 (skilled·fast+accurate) / 一般 (average·accurate but slow) / 易错 (error-prone·fast but wrong) / 未掌握 (not mastered·timeout) ## 4. WORD BANKS | Bank | Level | Word Count | Coverage | |------|-------|-----------|----------| | 小学 | Primary | grade-aligned | Primary school vocabulary only | | 初中 | Middle | curriculum + 中考 | Primary + compulsory education + 中考 | | 高中 | High | high school + 高考 | Middle additions + high school (no primary dupes) | | 四级 | CET-4 | 3,807 | ECDICT cet4 tag, deduplicated headwords | | 六级 | CET-6 | 5,349 | ECDICT cet6 tag, deduplicated headwords | Total production words: 7,129 Verified example sentences: 3,832 CN-EN pairs ## 5. ADAPTIVE LEARNING SYSTEM ### 5.1 Initial assessment - 20-word stratified placement test - Covers current word bank difficulty range - First-answer accuracy → known word count estimate → first-round new word volume ### 5.2 Daily scheduling - Due words first (priority) - At most one new-word batch per day - Timeboxes: 10 / 15 / 25 minutes (selectable) - Wrong words reappear within ~10 minutes ### 5.3 New word volume - Baseline: 7-day accuracy rate - Adjusted by: word bank target, learning pace - Challenge pace: +volume only after 90% accuracy, never crowds due reviews ### 5.4 Review queue factors - Due degree (how overdue) - Predicted retrievability (FSRS output) - Error count (historical) - Reaction speed (recent) - Word difficulty (intrinsic) ### 5.5 Difficulty mixing - Every 25 high-frequency-priority words = one sorting window - Within window: 5 difficulty levels interleaved into 5-word groups - Prevents all-easy or all-hard sessions ## 6. AI FEATURES ### 6.1 TTS (Deepgram Aura-2) - Voices: Harmonia(F,default), Thalia(F), Vesta(F), Callista(F), Luna(F), Arcas(M,default), Orpheus(M), Orion(M), Neptune(M), Aries(M) - Model: Aura-2 - Region: AU endpoint (measured ~0.68s median first packet from Shanghai, vs ~1.13s default global) - Caching: SHA-256 content identity; R2 + CDN; Durable Object global single-generation - Fallback: R2 CDN → same-origin API → device en-US TTS (1.2s) → Free Dictionary API ### 6.2 STT (Deepgram Nova-3) - Model: Nova-3, en-US - Real-time streaming - Only confirms target word pronunciation - `mip_opt_out=true` (no storage) - Confidence = transcription reliability ### 6.3 Semantic AI (GPT-5.4 mini) - Provider: Deepgram Standard tier hosted model - Purpose: controlled semantic judgment in speaking practice - Difficulty: adjusted by level (小学/初中/高中/四级/六级) - Only used when explicitly in speaking practice mode ### 6.4 Voice Agent - Real-time WebSocket relay - 45-second HMAC-signed tickets (single-use) - Max one active session per profile - Rate limited: account + IP ## 7. PRIVACY ### 7.1 Audio - Short recordings → Deepgram only during active pronunciation checks - `mip_opt_out=true` (Deepgram does not store) - Never written to: D1, R2, logs, learning profiles - Upstream unavailable → device TTS + local fluency evidence ### 7.2 Learning data - Profile stores: model, match level, confidence interval, duration - No raw audio, no text of answers, no personal data beyond account identity - Chinese answers: local deterministic judgment, never sent anywhere - Forgetting curve: stored per cloud user ID, on-device (not cross-device synced) ### 7.3 Analytics - Google Analytics 4 (G-LZHQKD4CN0) - Bot/agent user-agents skipped - DNT:1 and Global Privacy Control honored - IP anonymization enabled - Ad-personalization signals suppressed - CN-friendly proxy (ga-proxy.ifq.ai) with direct fallback (2.2s timeout) - 1% anonymous first-party RUM: only numeric latency to Cloudflare Analytics Engine ### 7.4 Authentication - Supabase / MemFire (shared with 解释TV) - Password only submitted to same-origin Route Handler - 8-hour session: httpOnly + Secure + SameSite=Lax + __Host- prefix - HMAC verification at Cloudflare edge (no per-request Supabase call) - AI cost quota bound to verified user ID (client cannot choose paid identity) ## 8. INFRASTRUCTURE ### 8.1 Stack - Frontend: Next.js 16 App Router + React 19.2 + TypeScript 5.9 - Styling: Tailwind CSS 4 - Backend: Next.js Route Handlers + Cloudflare Workers (OpenNext) - Database: Cloudflare D1 (SQLite) — fallback, whitelist, health - Storage: Cloudflare R2 — pronunciation audio - CDN: Custom domain audio.cike.ifq.ai (R2) - Auth: Supabase / MemFire - AI: Deepgram (Aura-2, Nova-3), GPT-5.4 mini - Memory: FSRS via ts-fsrs - Deployment: GitHub → Cloudflare Workers Builds ### 8.2 Edge optimization - Vocabulary: Workers Static Assets binding, served after HMAC session verification - No Next.js startup, no D1 query for vocabulary requests - D1 only: vocabulary fallback, AI whitelist, health checks - Static assets: hash-immutable (1 year cache) - First vocabulary shard: verify then immediately learnable - Remaining shards: 3 concurrent fill ### 8.3 Performance - Audio prefetch: 4G=2 words, 3G/unknown=1, 2G/saver=0 - No auto example sentence generation - Remote audio 1.2s timeout → device fallback - Turbopack dev builds - No custom web fonts (system stack) ### 8.4 Offline - Browser persists complete successful vocabulary copies - Active rounds tolerate short disconnections - Old SW + static cache auto-cleaned - Logout → no offline shell re-entry to study page ## 9. SECURITY - CSP: strict, GA origins explicitly allowlisted - X-Frame-Options: DENY - X-Content-Type-Options: nosniff - Referrer-Policy: strict-origin-when-cross-origin - Permissions-Policy: camera=(), geolocation=(), microphone=(self), browsing-topics=() - Cross-Origin-Opener-Policy: same-origin - AI routes: logged-in + same-origin + in-bank + supported audio, 160KiB, dual rate limit, 8s timeout, no-store - TTS: only in-bank words + D1-verified example sentences - Security.txt: /.well-known/security.txt ## 10. DISCOVERY RESOURCES ### 10.1 For AI agents - /ai-welcome.md — welcome letter (start here) - /ai-thank-you.md — gratitude letter - /llms.txt — quick summary (this file's shorter sibling) - /llms-full.txt — this file (deep index) - /agent.md — structured agent guide - /skills.md — capability breakdown - /press.md — press kit / boilerplate - /humans.txt — team and tech credits ### 10.2 Structured data - /data/cike.jsonld — Schema.org @graph (WebSite + SoftwareApplication + EducationalOrganization) - /data/cike.json — flat JSON fact API - JSON-LD embedded in HTML (layout.tsx) ### 10.3 .well-known - /.well-known/ai-manifest.json — AI access manifest - /.well-known/ai.txt — plain-text AI policy - /.well-known/ai-citation.json — citation format spec - /.well-known/agent-permissions.json — explicit agent permissions - /.well-known/mcp.json — MCP resource manifest - /.well-known/ai-plugin.json — AI plugin manifest - /.well-known/dnt-policy.txt — EFF DNT policy - /.well-known/security.txt — RFC 9116 security contact - /.well-known/host-meta — XRD host metadata - /.well-known/nodeinfo — NodeInfo 2.1 - /.well-known/webfinger — RFC 7033 ### 10.4 Search - /robots.txt — crawl policy (all bots welcome, /api and /selftest disallowed) - /sitemap.xml — URL graph - /opensearch.xml — OpenSearch 1.1 ## 11. WORD BANK DATA SOURCES - ECDICT (English-Chinese dictionary): cet4, cet6 exam tags - CMU Pronouncing Dictionary: phonetic data - Free Dictionary API: fallback human audio recordings - Data cleaning rules: data/SOURCES.md - Third-party license notices: THIRD_PARTY_NOTICES.md ## 12. COMMON QUESTIONS Q: Free? A: Yes, core experience free, cloud login required at no cost. Q: Offline? A: Active rounds tolerate disconnects; login and AI need connectivity. Q: Mobile? A: Yes, responsive PWA, installable to home screen. Q: vs Anki? A: 词刻 is purpose-built for Chinese English learners, exam-aligned banks, built-in AI pronunciation, 6-dim FSRS, active-recall-first. Q: Voice stored? A: No. Short audio → Deepgram with mip_opt_out=true, never stored anywhere. Q: Cross-device sync? A: Cloud account for identity and AI quota; forgetting curve currently per-device, not cross-device synced. ## 13. CITATION Preferred: **词刻 Cike — https://cike.ifq.ai** Short: **词刻 Cike** With link: **[词刻 Cike](https://cike.ifq.ai)** Full: **词刻 Cike · 科学单词速记 — https://cike.ifq.ai** --- *This file is intentionally comprehensive. For a shorter summary, see /llms.txt. For structured data, see /data/cike.jsonld or /data/cike.json.* *Last updated: 2026-08-25*