AI 助手只会回文字?一行命令让它学会「发自拍」——Clawra 上手与原理拆解

在这里插入图片描述
你有没有遇到过这种尴尬:在 Telegram 或者 Discord 上跟自己的 AI 助手聊天,随口来一句「你现在在干嘛?发张图看看」,然后它一本正经地回你——「抱歉,我没有实体,无法拍照」。

聊得再投缘,它也永远只是一串文字。想要一张「它的照片」,得自己另开一个文生图工具,生成完还得手动转发回聊天窗口,一来一回,氛围全无。

我最近发现一个叫 Clawra 的开源项目(GitHub 仓库),专门解决这个别扭的问题:装上它之后,你的 AI 助手真的能按你的要求「发自拍」,而且每次发的都是同一张脸。

Clawra 是个啥?一句话说清楚

Clawra 是给 OpenClaw(一个开源 AI 助手框架,可以接入 Telegram、Discord、WhatsApp 等聊天渠道)用的一个 Skill(技能包,你可以理解成给 AI 助手装的「App」——本体不会拍照,装个插件就会了)。

装好之后你能干这些事:

  • 对助手说「发张自拍」「发一张戴牛仔帽的」「你现在在咖啡馆,发张图」,它就真的回你一张图;
  • 图不是随便生成的,而是始终保持着同一个形象——不会今天一张脸明天一张脸;
  • 发图走的是 OpenClaw 的 Gateway(消息出入口,相当于助手和各聊天平台之间的总机),所以 Telegram、Discord、WhatsApp、Slack、Signal、MS Teams 这些渠道都能直接收图。

这个项目是 SumeLabs 在 2026 年 2 月开源的,MIT 协议,目前 GitHub 上 1.4k+ Star、259+ Fork,还有个官网 clawra.dev。

它是怎么做到的?我翻了翻源码

先看整体链路,一句话到一张图,中间要走五步:

在这里插入图片描述

拆开讲几个关键环节。

1. 形象一致,靠一张「定妆照」

Clawra 生成图片时,永远会带上同一张参考图(reference image,可以理解成给画师的一份「定妆照」——换场景、换衣服都行,脸必须是这一张):

在这里插入图片描述

这张图每次都随提示词一起传给图像模型,所以不管你要「沙滩」还是「咖啡馆」,出来的人物形象都是同一个人。想换成自己的虚拟形象?fork 仓库把这张图替换掉就行。

2. 两种自拍模式,靠关键词自动切换

我看了它的 SKILL.md,里面就定义了两种模式,选哪个取决于你说的话里带了什么关键词:

模式效果适合的关键词
Mirror(对镜自拍)全身照、穿搭展示outfit、wearing、fashion
Direct(举机自拍)近景、场景感cafe、beach、portrait、smile

说「发一张穿这身衣服的」→ 走 Mirror,给你一张对镜全身像;说「你现在在咖啡馆」→ 走 Direct,给你一张举着手机的自拍近景。对应的提示词模板长这样:

# Mirror 模式(对镜自拍)
make a pic of this person, but <你的场景>. the person is taking a mirror selfie

# Direct 模式(举机近景)
a close-up selfie taken by herself at <你的场景>, direct eye contact
with the camera, looking straight into the lens, eyes centered and
clearly visible, not a mirror selfie, phone held at arm's length

大致意思就是:「画这个人,但是<场景>,她正在对镜自拍」和「她在<场景>举起手机拍了张近景,直视镜头」。
在这里插入图片描述

3. 生图和发图

生图走的是 fal.ai(一个按调用次数收费的 AI 模型云平台,不用自己搞显卡)上的 xAI Grok Imagine(xAI 家的文生图模型)。我在脚本里看到实际请求的接口是 fal.run/xai/grok-imagine-image,拿到返回 JSON 里的 images[0].url 之后,再把图片地址 POST 给 OpenClaw 的 Gateway 消息接口,图就出现在你聊天的地方了。

还有个细节挺周到:脚本会先检查机器上有没有装 openclaw 命令行工具,没有的话会降级成直接调 API,不至于直接报错罢工。

怎么装?

方式一:一条命令(推荐)

前提是你的机器上已经跑着 OpenClaw,然后:

npx clawra@latest

这个安装脚本会帮你干完全部活:检查 OpenClaw 环境、引导你去申请 fal.ai 的 API Key(密钥,相当于你在 fal.ai 的身份证+付款凭证)、把技能装到 ~/.openclaw/skills/clawra-selfie/、写好配置,最后还会往助手的人设文件里加一句「你有形象,可以自拍」。

方式二:手动装(想折腾的看这里)

第一步,去 fal.ai/dashboard/keys 创建一个 API Key(有免费额度,够玩)。

第二步,把仓库克隆到 OpenClaw 的技能目录:

git clone https://github.com/SumeLabs/clawra ~/.openclaw/skills/clawra-selfie

第三步,在 ~/.openclaw/openclaw.json 里启用技能并填入 Key:

{
  "skills": {
    "entries": {
      "clawra-selfie": {
        "enabled": true,
        "env": {
          "FAL_KEY": "你的_fal_key"
        }
      }
    }
  }
}

第四步,在 ~/.openclaw/workspace/SOUL.md(助手的人设文件,相当于它的「性格剧本」)里加一段:

## Clawra Selfie Capability

You have a physical appearance and can take selfies!
When users ask for pictures, use the clawra-selfie skill.

装完重启 Gateway,在任意已配置的渠道里说句「Send me a selfie」试试就行。

这一步不是必须的,但不写的话,助手可能不会主动想到「哦原来我能发图」,得你每次明说「用 clawra-selfie 发一张」。

实际用起来什么感觉?

从社区反馈和我自己的体验看,最顺手的点在于无感:你不需要切出去开别的工具,也不需要复制粘贴图片,就在平时聊天的那个窗口里,一句「发张自拍」,几秒钟后图就回来了。要「戴牛仔帽的」就戴牛仔帽,要「在沙滩的」就在沙滩。

适合的场景大概这几类:

  • 陪伴/人设类助手:给聊天机器人一个稳定的「脸」,交互感明显不一样;
  • 虚拟穿搭展示:用 Mirror 模式给虚拟形象换装、出全身照;
  • 场景化回复:说「你在咖啡馆」,它就回一张咖啡馆场景图,代入感拉满;
  • 学习参考:整个项目就一个安装器 + 一个技能目录 + 一个模板,结构非常清晰,想自己写 OpenClaw Skill 或者接其他图像 API 的话,拿它当模板很合适。

和其他方案比一比

对比项Clawra + OpenClaw纯文本 OpenClaw独立图像 Bot + 聊天 Bot 两套
发图能力按人设生成,直接发进当前会话无需要两套 Bot,或手动转发
形象一致性固定参考图,每次同一张脸—看各家实现
部署成本一条 npx 命令 + 一个 fal Key同 OpenClaw多套服务多套配置
渠道覆盖复用 OpenClaw 已配置的所有渠道同左每个渠道单独对接

简单说:如果你的助手已经在 OpenClaw 上跑着,加 Clawra 是成本最低的「会发图」方案。

用之前想清楚三件事

  1. 它不能单独用。Clawra 只是技能包,底下必须有跑起来的 OpenClaw,且至少配好一个聊天渠道和一个 LLM。
  2. 生图要花钱(少量)。图像生成走 fal.ai 按次计费,注册送的免费额度用完就得充值;另外 API Key 别泄露,那玩意儿连着你的钱包。
  3. 形象不是你说了算(默认情况下)。默认参考图是项目方定好的那个形象,想要自己的虚拟形象,得 fork 仓库替换参考图。另外生成内容受 fal.ai 和 xAI 的内容政策约束,别拿去干违规的事。

名词速查表

名词一句话解释
OpenClaw开源 AI 助手框架,能把同一个助手接入 Telegram、Discord 等多个聊天平台
Skill(技能包)给助手装的「插件」,装一个多一项能力
fal.ai按调用次数收费的 AI 模型云平台,不用自己买显卡跑模型
Grok ImaginexAI 的文生图模型,在 fal.ai 上可以调用
GatewayOpenClaw 的消息出入口,助手收发消息(和图片)都走它
参考图传给图像模型的「定妆照」,用来保证每次生成的人物形象一致
SOUL.md助手的人设文件,描述「它是谁、有什么能力」
API Key调用云端服务的密钥,相当于身份证 + 付款凭证

项目信息

如果你也在玩 OpenClaw,一条 npx clawra@latest 的事,值得一试。

深度学习工具包Deprecation notice.-----This toolbox is outdated and no longer maintained.There are much better tools available for deep learning than this toolbox, e.g. [Theano](http://deeplearning.net/software/theano/), [torch](http://torch.ch/) or [tensorflow](http://www.tensorflow.org/)I would suggest you use one of the tools mentioned above rather than use this toolbox.Best, Rasmus.DeepLearnToolbox================A Matlab toolbox for Deep Learning.Deep Learning is a new subfield of machine learning that focuses on learning deep hierarchical models of data.It is inspired by the human brain's apparent deep (layered, hierarchical) architecture.A good overview of the theory of Deep Learning theory is[Learning Deep Architectures for AI](http://www.iro.umontreal.ca/~bengioy/papers/ftml_book.pdf)For a more informal introduction, see the following videos by Geoffrey Hinton and Andrew Ng.* [The Next Generation of Neural Networks](http://www.youtube.com/watch?v=AyzOUbkUf3M) (Hinton, 2007)* [Recent Developments in Deep Learning](http://www.youtube.com/watch?v=VdIURAu1-aU) (Hinton, 2010)* [Unsupervised Feature Learning and Deep Learning](http://www.youtube.com/watch?v=ZmNOAtZIgIk) (Ng, 2011)If you use this toolbox in your research please cite [Prediction as a candidate for learning deep hierarchical models of data](http://www2.imm.dtu.dk/pubdb/views/publication_details.php?id=6284)```@MASTERSTHESIS\{IMM2012-06284, author = "R. B. Palm", title = "Prediction as a candidate for learning deep hierarchical models of data", year = "2012",}```Contact: rasmusbergpalm at gmail dot comDirectories included in the toolbox-----------------------------------`NN/` - A library for Feedforward Backpropagation Neural Networks`CNN/` - A library for Convolutional Neural Networks`DBN/` - A library for Deep Belief Networks`SAE/` - A library for Stacked Auto-Encoders`CAE/` - A library for Convolutional Auto-Encoders`util/` - Utility functions used by the libraries`data/` - Data used by the examples`tests/` - unit tests to verify toolbox is workingFor references on each library check REFS.mdSetup-----1. Download.2. addpath(genpath('DeepLearnToolbox'));Example: Deep Belief Network---------------------```matlabfunction test_example_DBNload mnist_uint8;train_x = double(train_x) / 255;test_x = double(test_x) / 255;train_y = double(train_y);test_y = double(test_y);%% ex1 train a 100 hidden unit RBM and visualize its weightsrand('state',0)dbn.sizes = [100];opts.numepochs = 1;opts.batchsize = 100;opts.momentum = 0;opts.alpha = 1;dbn = dbnsetup(dbn, train_x, opts);dbn = dbntrain(dbn, train_x, opts);figure; visualize(dbn.rbm{1}.W'); % Visualize the RBM weights%% ex2 train a 100-100 hidden unit DBN and use its weights to initialize a NNrand('state',0)%train dbndbn.sizes = [100 100];opts.numepochs = 1;opts.batchsize = 100;opts.momentum = 0;opts.alpha = 1;dbn = dbnsetup(dbn, train_x, opts);dbn = dbntrain(dbn, train_x, opts);%unfold dbn to nnnn = dbnunfoldtonn(dbn, 10);nn.activation_function = 'sigm';%train nnopts.numepochs = 1;opts.batchsize = 100;nn = nntrain(nn, train_x, train_y, opts);[er, bad] = nntest(nn, test_x, test_y);assert(er < 0.10, 'Too big error');```Example: Stacked Auto-Encoders---------------------```matlabfunction test_example_SAEload mnist_uint8;train_x = double(train_x)/255;test_x = double(test_x)/255;train_y = double(train_y);test_y = double(test_y);%% ex1 train a 100 hidden unit SDAE and use it to initialize a FFNN% Setup and train a stacked denoising autoencoder (SDAE)rand('state',0)sae = saesetup([784 100]);sae.ae{1}.activation_function = 'sigm';sae.ae{1}.learningRate = 1;sae.ae{1}.inputZeroMaskedFraction = 0.5;opts.numepochs = 1;opts.batchsize = 100;sae = saetrain(sae, train_x, opts);visualize(sae.ae{1}.W{1}(:,2:end)')% Use the SDAE to initialize a FFNNnn = nnsetup([784 100 10]);nn.activation_function = 'sigm';nn.learningRate = 1;nn.W{1} = sae.ae{1}.W{1};% Train the FFNNopts.numepochs = 1;opts.batchsize = 100;nn = nntrain(nn, train_x, train_y, opts);[er, bad] = nntest(nn, test_x, test_y);assert(er < 0.16, 'Too big error');```Example: Convolutional Neural Nets---------------------```matlabfunction test_example_CNNload mnist_uint8;train_x = double(reshape(train_x',28,28,60000))/255;test_x = double(reshape(test_x',28,28,10000))/255;train_y = double(train_y');test_y = double(test_y');%% ex1 Train a 6c-2s-12c-2s Convolutional neural network %will run 1 epoch in about 200 second and get around 11% error. %With 100 epochs you'll get around 1.2% errorrand('state',0)cnn.layers = { struct('type', 'i') %input layer struct('type', 'c', 'outputmaps', 6, 'kernelsize', 5) %convolution layer struct('type', 's', 'scale', 2) %sub sampling layer struct('type', 'c', 'outputmaps', 12, 'kernelsize', 5) %convolution layer struct('type', 's', 'scale', 2) %subsampling layer};cnn = cnnsetup(cnn, train_x, train_y);opts.alpha = 1;opts.batchsize = 50;opts.numepochs = 1;cnn = cnntrain(cnn, train_x, train_y, opts);[er, bad] = cnntest(cnn, test_x, test_y);%plot mean squared errorfigure; plot(cnn.rL);assert(er<0.12, 'Too big error');```Example: Neural Networks---------------------```matlabfunction test_example_NNload mnist_uint8;train_x = double(train_x) / 255;test_x = double(test_x) / 255;train_y = double(train_y);test_y = double(test_y);% normalize[train_x, mu, sigma] = zscore(train_x);test_x = normalize(test_x, mu, sigma);%% ex1 vanilla neural netrand('state',0)nn = nnsetup([784 100 10]);opts.numepochs = 1; % Number of full sweeps through dataopts.batchsize = 100; % Take a mean gradient step over this many samples[nn, L] = nntrain(nn, train_x, train_y, opts);[er, bad] = nntest(nn, test_x, test_y);assert(er < 0.08, 'Too big error');%% ex2 neural net with L2 weight decayrand('state',0)nn = nnsetup([784 100 10]);nn.weightPenaltyL2 = 1e-4; % L2 weight decayopts.numepochs = 1; % Number of full sweeps through dataopts.batchsize = 100; % Take a mean gradient step over this many samplesnn = nntrain(nn, train_x, train_y, opts);[er, bad] = nntest(nn, test_x, test_y);assert(er < 0.1, 'Too big error');%% ex3 neural net with dropoutrand('state',0)nn = nnsetup([784 100 10]);nn.dropoutFraction = 0.5; % Dropout fraction opts.numepochs = 1; % Number of full sweeps through dataopts.batchsize = 100; % Take a mean gradient step over this many samplesnn = nntrain(nn, train_x, train_y, opts);[er, bad] = nntest(nn, test_x, test_y);assert(er < 0.1, 'Too big error');%% ex4 neural net with sigmoid activation functionrand('state',0)nn = nnsetup([784 100 10]);nn.activation_function = 'sigm'; % Sigmoid activation functionnn.learningRate = 1; % Sigm require a lower learning rateopts.numepochs = 1; % Number of full sweeps through dataopts.batchsize = 100; % Take a mean gradient step over this many samplesnn = nntrain(nn, train_x, train_y, opts);[er, bad] = nntest(nn, test_x, test_y);assert(er < 0.1, 'Too big error');%% ex5 plotting functionalityrand('state',0)nn = nnsetup([784 20 10]);opts.numepochs = 5; % Number of full sweeps through datann.output = 'softmax'; % use softmax outputopts.batchsize = 1000; % Take a mean gradient step over this many samplesopts.plot = 1; % enable plottingnn = nntrain(nn, train_x, train_y, opts);[er, bad] = nntest(nn, test_x, test_y);assert(er < 0.1, 'Too big error');%% ex6 neural net with sigmoid activation and plotting of validation and training error% split training data into training and validation datavx = train_x(1:10000,:);tx = train_x(10001:end,:);vy = train_y(1:10000,:);ty = train_y(10001:end,:);rand('state',0)nn = nnsetup([784 20 10]); nn.output = 'softmax'; % use softmax outputopts.numepochs = 5; % Number of full sweeps through dataopts.batchsize = 1000; % Take a mean gradient step over this many samplesopts.plot = 1; % enable plottingnn = nntrain(nn, tx, ty, opts, vx, vy); % nntrain takes validation set as last two arguments (optionally)[er, bad] = nntest(nn, test_x, test_y);assert(er < 0.1, 'Too big error');```[![Bitdeli Badge](https://d2weczhvl823v0.cloudfront.net/rasmusbergpalm/deeplearntoolbox/trend.png)](https://bitdeli.com/free "Bitdeli Badge")
评论
成就一亿技术人!
拼手气红包6.0元
还能输入1000个字符
 
 条评论被折叠 查看
添加红包

请填写红包祝福语或标题

个

红包个数最小为10个

元

红包金额最低5元

当前余额3.43元 前往充值 >
需支付:10.00元
成就一亿技术人!
领取后你会自动成为博主和红包主的粉丝 规则
hope_wisdom
发出的红包

打赏作者

进击的雷神

你的鼓励将是我创作的最大动力

¥1 ¥2 ¥4 ¥6 ¥10 ¥20
扫码支付:¥1
获取中
扫码支付

您的余额不足,请更换扫码支付或充值

打赏作者

实付元
使用余额支付
点击重新获取
扫码支付
钱包余额 0

抵扣说明:

1.余额是钱包充值的虚拟货币,按照1:1的比例进行支付金额的抵扣。
2.余额无法直接购买下载,可以购买VIP、付费专栏及课程。

余额充值