account-research
Research a company or person and get actionable sales intel. Works standalone with web search, supercharged when you connect enrichment tools or your CRM.…
\"Implement LDA topic modeling to discover latent topics in document collections. Use this skill when the user needs to extract topics from a text corpus, categorize documents by theme, or explore thematic structure — even if they say 'what are the main topics', 'topic
$ npx -y skills add charlieviettq/awesome-agent-skill --skill algo-nlp-lda --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/algo-nlp-ldaContext preview
The summary Claude sees to decide when to auto-load this skill.
\"Implement LDA topic modeling to discover latent topics in document collections. Use this skill when the user needs to extract topics from a text corpus, categorize documents by theme, or explore thematic structure — even if they say 'what are the main topics', 'topic
name: "\"algo-nlp-lda\"" description: "\"Implement LDA topic modeling to discover latent topics in document collections. Use this skill when the user needs to extract topics from a text corpus, categorize documents by theme, or explore thematic structure — even if they say 'what are the main topics', 'topic extraction', or 'document clustering by theme'.\"." allowed-tools: Read, Glob, Grep
Latent Dirichlet Allocation models each document as a mixture of topics and each topic as a distribution over words. Discovers K latent topics from a corpus without supervision. Uses Gibbs sampling or variational inference. Complexity: O(N × K × iterations) where N = total word tokens.
**Trigger conditions:**
**When NOT to use:**
IRON LAW: The Number of Topics K Must Be Chosen, Not Discovered LDA does NOT tell you how many topics exist. K is a hyperparameter. Too few topics: overly broad, mixed themes. Too many: fragmented, redundant topics. Use coherence score (C_v) to compare K values, but the final choice requires human judgment on topic interpretability.
Preprocess: tokenize, remove stop words, apply lemmatization. Build document-term matrix. Filter: remove terms appearing in <5 or >50% of documents. **Gate:** Clean DTM, vocabulary size reasonable (1K-50K terms).
1. Choose K (start with √(N/2), try range K=5,10,15,20,...) 2. Set hyperparameters: α = 50/K (document-topic density), β = 0.01 (topic-word density) 3. Run LDA (Gibbs sampling: 1000+ iterations, or variational inference) 4. Extract: topic-word distributions (top 10-20 words per topic) and document-topic distributions
Evaluate: topic coherence (C_v score, higher is better), manual inspection of top words per topic, check for "junk" topics (mixed incoherent words). **Gate:** Coherence score acceptable, topics are humanly interpretable.
Return topics with top words and document assignments.
{
"topics": [{"id": 0, "label": "finance", "top_words": ["revenue", "profit", "quarter", "growth"], "coherence": 0.55}],
"doc_topics": [{"doc_id": "d1", "dominant_topic": 0, "topic_distribution": [0.7, 0.1, 0.2]}],
"metadata": {"K": 10, "coherence_avg": 0.48, "documents": 5000, "vocabulary": 8000}
}**Input:** 1000 news articles, K=5 **Expected:** Topics like: {politics, sports, technology, business, entertainment} with coherent top words per topic.
| Input | Expected | Why | |-------|----------|-----| | Very short documents | Poor topic assignment | Too few words for reliable mixture estimation | | Homogeneous corpus | 1-2 topics dominate | All documents are similar, limited topic diversity | | K=1 | Single topic = corpus vocabulary | Degenerate case, no discrimination |
Curated skill pack for LLM agents in engineer and science workflow (Cursor & Claude ready).
Research a company or person and get actionable sales intel. Works standalone with web search, supercharged when you connect enrichment tools or your CRM.…
Evaluate LLM agents and tool-using workflows—task success, tool accuracy, latency/cost, safety, and regression suites. Use when shipping agent features,…
Design agent tools and CLI surfaces—schemas, naming, errors, idempotency, and discoverability for LLM callers. Use when defining tools for agents, SDKs, or…
\"Implement and select ad bidding strategies from manual CPC to automated target-CPA and target-ROAS. Use this skill when the user needs to choose a bidding…
\"Optimize advertising budget allocation across campaigns using marginal returns analysis. Use this skill when the user needs to distribute budget across…
\"Build CTR prediction models for estimating ad click-through rates from features. Use this skill when the user needs to predict click probability, build an ad…