<<返回 西安电子科技大学—学生翻译实践成果展
以下内容为西安电子科技大学学生最近一个月内的翻译实践成果

——中石油新疆销售有限公司阿克苏分公司红旗坡加油(气)站的蝶变重生之路 戈壁延绵,国道蜿蜒。在314国道阿克苏段的茫茫戈壁滩上,阿克苏销售公司红旗坡加油(气)站静静伫立。数十年栉风沐雨,这座站见证着南疆物流经济的起落更迭,镌刻着能源行业转型变革的时代年轮。这里曾是声名赫赫的万吨功勋站,车流如织、荣光满身;也曾因路网变迁、能源迭代陷入低谷,沉寂落寞、步履维艰。 从巅峰滑落低谷,于绝境逆势重生,一座基层站点的蝶变,正是阿克苏销售公司直面市场变局、主动破局转型的生动缩影。二十载党龄的老站长初心如磐、默默坚守,二十年工龄的外包骨干深耕一线、无私奉献,一群平凡的石油人,在风沙戈壁间扛压力、闯新路、拼突围,用不服输、不放弃、不止步的铁军精神,让褪色的功勋站牌重焕光彩,让红色战旗始终在红旗坡上空高高飘扬。 往昔峥嵘:万千车流里的功勋荣光 本世纪第一个10年,是南疆经济飞速腾飞的黄金岁月。314国道作为贯通南北疆的交通大动脉,承载着全域绝大多数货运、商贸车流。彼时的红旗坡加油站,占据国道黄金区位,是往来车辆必经的补给驿站。 那些年的红旗坡站,从晨光微熹到星夜沉沉,始终车水马龙、人声鼎沸。加油岛上,加油枪起落不停,员工们身着工装,穿梭在车流之间,引导车辆、扫码加油、答疑解惑,动作娴熟、步履匆匆。长长的货车队伍沿着国道有序排开,引擎轰鸣声、员工叮嘱声、车辆鸣笛声交织在一起,奏响着繁华鼎沸的乐章。依托源源不断的车流红利,站点成品油年销量一度突破万吨,是一座不折不扣的“功勋加油站”。 亮眼的经营业绩背后,是站点持之以恒的队伍建设与温度服务。不同于单纯追求销量的站点,红旗坡站始终把员工凝聚、班组建设放在首位,用心用情打造有温度、有力量的“职工之家”。站内管理制度暖心规范,师徒传帮带薪火相传,员工互帮互助、凝心聚力,形成了团结奋进、实干担当的优良班风。凭借扎实的班组建设成果和浓厚的人文氛围,让红旗坡加油站成为新疆公司在阿克苏销售公司召开“职工之家”建设现场会的现场地,将基层工会组织建设经验在全疆推广。红旗坡站不仅是国道旁显赫的万吨站,更是新疆销售基层队伍建设的样板标杆。一代代员工在这里扎根成长、实干奉献,把青春和汗水挥洒在戈壁国道,让这座小小的加油站,成为314国道上最亮眼的红色坐标。 风沙遇冷:时代变局下的沉寂阵痛 时代的滚滚浪潮,从不偏爱一成不变的坚守。就在红旗坡站大步发展、屡创佳绩之时,区域交通格局与能源结构的双重变革,让这座功勋站点骤然跌入寒冬。 G3012吐和高速公路全线通车后,南北疆长途货运、客运车辆纷纷改道高速,通行效率更高、路况更优的高速路网,极大地分流了314国道的滚滚车流。仿佛是一夜之间,曾经川流不息的国道骤然冷清,过往重卡、商贸车辆寥寥无几。几乎与此同时,新能源汽车快速普及、清洁能源加速替代,传统成品油市场持续承压,双重冲击之下,红旗坡站的经营状况急转直下。 繁华落幕,沉寂来袭。曾经终日忙碌的加油岛变得空旷冷清,频繁起落的加油枪渐渐静默,站内日均车流断崖式下跌,销售量持续走低。最艰难的时期,站点年成品油销量锐减至2000吨左右,不足巅峰时期的五分之一。 那是创造过历史、见证过辉煌的红旗坡加油站员工们最煎熬的一段日子。 也是现任站经理胡永斌,最为感慨的岁月。 2022年1月,这个土生土长的南疆汉子受命于危难之中,走马上任,成为这座日渐沉寂的小加油站第7任带头人。作为有着20年党龄的老党员、站点负责人,他作为旁观者,曾亲眼见证了红旗坡的鼎盛繁华,现如今,他要亲身承受市场寒冬的沉重压力。刚到站就职的那段时间,每天清晨,他习惯性地站在站区路口眺望,往日车水马龙的景象不复存在,空旷的国道、零星的过往车辆,让他心里满是落差与焦灼。站内员工士气低落,有人私下议论“这个站不行了”,有人萌生调岗、离职的想法,迷茫与不安笼罩着整个站点。 一次夜里值班,胡永斌看着空荡荡的加油场地,指尖摩挲着加油站一块块荣誉奖牌,斑驳的灯光落在沧桑的牌匾上,格外刺眼。“老站的荣光,不能毁在我们这一代人手里!”他在心里默默立下誓言:市场可以降温,但红旗坡的奋斗温度绝不能降;销量可以下滑,但石油人的精神脊梁绝不能弯! 在全员低落的氛围中,胡永斌默默扛起所有压力,他从不向员工抱怨困境,也不向上级推诿难题,只是日复一日坚守在岗,默默观察市场、梳理数据、寻找出路,带领全站员工在沉寂的寒冬里,苦苦等待破局的微光。 逆风破局:默默坚守中的实干突围 危局之中,唯勇者进;困境之下,唯实干兴。面对路网变革、能源迭代的双重考验,阿克苏销售公司党委“一班人”精准研判市场趋势,摒弃“守摊子、等回暖”的消极心态,抢抓清洁能源发展机遇,果断为红旗坡站量身定制转型方案,决定新增LNG加气业务,推动传统成品油站点向油气协同的综合能源服务站转型,为老牌功勋站注入全新生命力。 转型之路,从零起步、步步维艰。场地改造、设备安装、资质报批、人员培训、客户开拓,每一项工作都是全新挑战。在这场逆风突围的攻坚战中,站经理胡永斌带着核算员何海龙,这位入职中石油超过20年的老员工,成为站点转型最坚实的力量,以平凡坚守书写着不凡担当。 老站长的“戈壁拓客路”,以初心换人心。不善言辞的胡永斌,没有惊天动地的誓言,只有脚踏实地的行动。为打开LNG市场销路,他主动走出去,从“坐商”变“行商”,开启了漫漫拓客之路。盛夏的南疆戈壁,地表温度高达四五十摄氏度,热浪翻滚、风沙扑面,他日复一日走访周边物流园、货运车队、工矿企业,用脚步丈量着一名石油老兵思维和视野的广度。 有一次,他连续三天走访一家大型物流车队,负责人始终心存顾虑,担心新站点设备不稳定、加气服务无保障,一次次婉拒合作。连用户单位的人都看不下去了,私下里劝胡永斌放弃,可他偏不服输。第四天清晨,他早早等候在物流园门口,跟着出车司机逐一沟通,耐心讲解中国石油的“品牌故事”和保供能力,将我们的服务、价格、管理等优势,主动承诺优先保障车队加气、提供24小时便民服务。真诚终能打动人心,车队负责人被他的执着打动,率先与站点达成长期合作协议,成为红旗坡站LNG业务的第一批固定客户。 管理经验虽然丰富,但是LNG业务对胡永斌这个资深站经理来说,确实是一个全新的甚至是空白的领域,面临着不小的困难和挑战。“不会就学,不懂就问,反正必须得把新业务拿下来!”LNG设备调试期间,他全程驻守现场,年过五十的他,跟着年轻技术人员熬夜钻研设备原理、熟记操作规范、排查安全隐患,厚厚的操作规程被他翻得卷边,密密麻麻的笔记写了不知多少。他说:“我是站经理,更是老党员,新业务我先学、难题我先上,才能带着大家干得好!” 老骨干的“日夜坚守岗”,以匠心守初心。核算员何海龙,作为业务外包员工的优秀代表,二十年扎根一线,把最好的青春献给了戈壁石油站点。LNG业务落地后,零基础、新技能的难题摆在眼前,年近五十的他,主动扛起攻坚重任。为快速掌握加气技能,他利用休息时间加班加点,逐字逐句逐字逐句学习安全操作规程,反复演练加气操作、应急处置、设备巡检流程。别人休息时,他在模拟操作;夜深人静时,他在复盘细节。短短一个月,他从“门外汉”变成了“行家里手”,熟练掌握LNG基本的接卸、加注业务。 长途货车大多夜间通行、凌晨加气,为抢抓客流,何海龙常常主动上全站最难、最熬人的夜班。今年早春的一个深夜,戈壁狂风大作、气温骤降,一辆重型半挂车缓缓驶入站区。彼时已是凌晨三点,寒风刺骨、沙尘漫天,何海龙顶着大风上前引导车辆、规范加气。加注过程中,他发现司机面色疲惫、手脚冰凉,加气结束后,他主动递上一杯滚烫的热水,又跑前跑后地帮司机泡上一桶热气腾腾的泡面,贴心叮嘱司机注意行车安全。 “跑遍西北这么多加气站,真就没见过这么贴心的员工!都知道LNG大车加一次液能跑好久,下次来不知道得是什么时候了,可在这个站,根本没有那种‘一锤子买卖’的感觉!”司机的由衷夸赞,是对何海龙二十年坚守最好的认可。日复一日,他用细微暖心的服务积攒客户、留住口碑,越来越多货车司机宁愿绕道,也要来红旗坡站加液休憩。 浴火重生:老站新程的可期未来 一人带头冲锋,全员聚力攻坚。在胡永斌、何海龙的示范带动下,红旗坡站全体员工拧成一股绳、合成一股劲,开启了全方位的突围蜕变。站内全员化身销售员、服务员、宣传员,线上搭建司机沟通群,实时推送价格优惠、路况提示、便民服务;线下分片走访拓客,深耕周边市场、稳固存量客源、拓展增量用户。同时,站点持续优化服务品质,打造暖心驿站,用专业、贴心、高效的服务重塑站点口碑。 汗水浇灌硕果,奋斗终获回响。历经艰难的蜕变重生、业务转型,红旗坡站彻底走出销量低谷,实现了华丽重生。如今,该站LNG日均销量稳定在30吨左右,同时带动油品和非油业务销售。刚刚过去的这个7月,红旗坡加油站油品销售同比增长5%,非油收入同比增长近50%,非油毛利更是劲增63%,油气非业务真正实现了互补共进、稳步攀升。截至7月末,该站已经完成全年燃气业务指标的93.48%,超时间进度35.15个百分点。按照当前喜人的经营态势,全站年度油气当量有望再度突破万吨大关! 从年销万吨的巅峰落幕,到销量“膝斩”的低谷蛰伏,再到油气并举、重回巅峰的强势突围,红旗坡站的涅槃重生,绝非偶然,更非简单的销量回升,而是阿克苏销售公司党委高瞻远瞩、审视全局、精准研判车用能源变革大势的科学成果,是公司立足行业转型浪潮、高位推动基层站点提质增效、布局综合能源服务赛道的生动实践。面对站点发展困局,公司没有放任老站沉寂、被动应对市场变局,而是跳出单点经营思维,立足企业整体发展格局,精准把脉路网变革、能源替代带来的机遇与挑战,科学制定转型方案、统筹调配资源、全力攻坚项目落地,为红旗坡站脱困突围指明了方向、筑牢了根基。自上而下的科学部署、精准赋能,以及红旗坡站全体员工凝心聚力、百折不挠的拼搏实干,成就了这场深刻的自我革新与精神重塑。如今,轰鸣的加液枪和着加油枪欢快歌唱,曾经的“单音调”升华成了如今的“协奏曲”,冷清的站区重归热闹繁忙,曾几何时黯淡褪色的金字招牌,在戈壁阳光下重新熠熠生辉。 时代的浪潮滚滚向前,能源的变革正逢其时。当下,能源结构持续迭代,市场竞争日趋激烈,传统石油销售行业依旧面临重重挑战。但在红旗坡站,在阿克苏销售公司,变的是经营赛道、服务模式、业态格局,不变的是自强不息、迎难而上的奋斗底色,是敢闯敢拼、逆境突围的铁军精神。 戈壁辽阔,初心不改;红旗猎猎,使命不息。今天的胡永斌、何海龙和全站员工依旧坚守在314国道旁,以岗为家、以业为荣,在平凡的岗位上默默耕耘、接续奋斗。这座历经沉浮的老牌功勋站,正以全新的姿态、昂扬的斗志,在能源转型的新征程上破浪前行,让鲜红的奋斗旗帜,永远高高飘扬在红旗坡的戈壁之上!

2026-08-18 李宇轩 化工 中-英

比亚迪真面目竟是“深圳富士康”? 比亚迪大利空? 刚刚发布的半年报显示,这家市值8104.76亿的中国新能源汽车龙头,今年上半年只赚了11.74亿。 说“只赚”,是相比自主品牌头部车企的长城、吉利等等对手,它们的上半年利润分别是35.29亿、24亿,并且这两家市值远不及比亚迪。 这业绩好像是有点… 而且,从比亚迪集团全局来看,手机代工业务超过汽车,成为最大的营收版块。 难道比亚迪真面目竟是“深圳富士康”? 离谱吗? 这还不是最离谱的。汽车业务不挣钱的背后,竟然还有产销增幅过50%的业绩。 看不懂了。 车卖得越好越不挣钱,新能源一哥比亚迪,这半年发生了什么? “深圳富士康”? 比亚迪上半年财报透露出的经营状况,用喜忧参半来形容再合适不过。 喜从何来? 先看营收,今年上半年,比亚迪实现总营收908.85亿元人民币,同比上期增长50.22%。 与此同时,在营收构成上,也出现了一些有意思的变化。 以往比亚迪最大业务板块——汽车业务,上半年实现营收391.57亿元人民币,同比上期增加22.09%。 营收上涨当然是车卖得更红火了。根据财报,比亚迪上半年汽车销量累计达到24.67万辆,同比上期增加55.51%。 其中,新能源车销量15.46万辆,同比增幅达到154.76%,占到总体销量的62.67%。 虽然有所增长,但是作为“车企”的比亚迪,汽车板块失掉了集团内“C位”。汽车业务营收从去年同期的53.01%下降到43.08%,降幅将近10个百分点。 瓜分汽车业务份额的,是手机部件及组装业务和二次充电电池及光伏业务。 根据财报,上半年比亚迪的手机部件及组装业务贡献了431.32亿元人民币营收,同比上期大增84.48%。 在营收占比上,由原来的38.64%增加到47.46%,手机部件及组装业务成为营收贡献最大的板块。 除此之外,二次充电电池及光伏业务也有所增加,上半年实现营收82.87亿元人民币,同比增加72.97%。 营收占比涨幅不大,由去年同期的7.92增加到9.12%。、 没想到吧? 当大家都以为比亚迪是电池起家的自主新能源汽车领军者时,实际上手机代工业务才是人家最大的营收来源。 但营收和销量双双增长,钱却没赚多少。 财报显示,上半年比亚迪实现归属上市公司股东净利润11.74亿元人民币,同比上期减少29.41%。 如果除去政府补助等非经常性损益,归属上市公司股东净利润就只剩下3.67亿元人民币,同比大跌60%。 问题到底出在哪儿了? 卖车不好赚了? 先看毛利率。 上半年,比亚迪毛利率12.76%,较去年同期下滑6.84个百分点。 分业务来看,汽车业务上半年毛利率19.53%,虽然高于整体水平,但是较上年同期也下滑4.4个百分点。 而对营收贡献最大的手机组装代工业务,上半年毛利率只有7%,较上年同期下滑6.64个百分点。 结合比亚迪业务结构的变动可以看出,利润丰厚的汽车业务占比下降,只能赚点“辛苦钱”的手机代工却成了最大的营收来源。 除了业务结构的变动,还有一个变化值得注意,不管是哪一块,比亚迪上半年的毛利率都在下跌。 引起这一变动的,是电子产品、新能源相关的原材料价格上涨所致。 公开消息显示,最近一年来,国际铜价、铝价涨幅约50%。 比亚迪电池生产重要原材料碳酸锂全年上升幅度甚至达90%。 六氟磷酸锂从20年7月的不足7万元/吨上涨至21年7月的40万元/吨,期间涨幅近470%。 △资料来自钢联数据 而伊维经济研究院测算,磷酸铁锂正极材料上涨对磷酸铁锂电池影响比例为3.77%,主流三元正极材料上涨对三元锂电池影响相对较高,影响比例超过10%。 除此之外,手机组装代工行业重要的成本构成——人工成本的增加也是压缩比亚迪利润空间的重要因素。 财报显示,上半年比亚迪用于短期薪酬支出较上期增加124.08亿元人民币,其中劳务派遣费用支出增加22.13亿元。 比亚迪挣“辛苦钱” 加量不加价、增收不增利,比亚迪这个中国新能源龙头,到底咋回事? 长期以来,比亚迪集团的汽车业务贡献营收一半、利润70%左右,手机代工业务营收尽管过4成,但利润徘徊在20%左右。 手机代工企业虽有硬核实力,但赚的却是辛苦钱,比亚迪也不例外,其集团绝对主力一直都是汽车业务。 但如今毛利低的手机板块业绩占比提升,毛利高的汽车板块业绩占比下降,这本来和企业的逐利本性是不相符的。 比亚迪自己给出的主要原因是原材料成本增加。 此话不假。确切的说,是整个制造业上游成本增加,大家的日子都不太好过。 数据显示,7月工业企业利润同比增速已降至八个月新低,尤其中下游企业盈利空间不断受到挤压。 △资料来源:招商银行研究院、Macrobond 一个直观证据是,体现经济走势强弱的制造业采购经理人指数,7月官方数据降至50.4,创17个月最低。 简单理解,超过50则扩张,低于50则萎缩。 制造业挣扎在荣枯线附近,是普遍现象,比亚迪自然不能幸免。 对于比亚迪来说,它的上游包括采矿、原材料制造(钢铁、橡胶等等)等等,7月这两个行业利润同比分别增长2.03倍、50.9%。 当然,还有最重要的,锂业。这也是比亚迪区别于其他车企最明显的“新能源”标签的根本。 自去年第四季以来,锂原料成本大涨,上面已经列举。 这也是为何同是制造业巨头的长城、吉利,在原材料价格上涨情况下,却能守住利润基本盘。 因为,它们都是燃油车大户,和比亚迪根本不在一个生态位。 上游成本普遍增加,且动力电池成本又来了一波精准打击,比亚迪最成功的的“招牌”成了最严重的“灾区”,产销大涨也挽救不了。 所以,这半年比亚迪看似车卖得风生水起,市占率达16%,把新能源一哥的位置坐得更稳,实际却不得不靠手机代工那份“辛苦钱”弥补利润下滑。 都看到做大哥风光,谁了解大哥背后的辛苦… 那么,接下来比亚迪会怎么办? 根据公开数据,销量上,比亚迪已经超越特斯拉成国产新能源第一,寄予厚望的旗舰车型汉,累计销售10万辆,成为市场强力车型。 另外,比亚迪一直在憋的一个“大招”是一系列DM-i混动车型。 上半年中,6月销售DM车型20100辆,同比增长536.7%;7月销售DM车型25061辆,同比增长650.6%。 在国内日系“两田”混动夹击下,月销过两万已经是竞争力的证明。 况且这还是DM-i混动车型刚刚上市,产能尚未完全释放的情况。 所以只要守住市场份额,捱过这波原材料成本波动,总能迎来希望。 此外,在品牌上,比亚迪计划年底将发布高端品牌,首款车型预计明年亮相。 通过拉高品牌调性,增加溢价,的确是提高利润率的好办法。 只是,冲击高端绝非有雄心就能实现,“手机不挣钱”的小米雷军,可能更知其味。 对了,今年早些时候王传福还苦口婆心劝雷军造车需谨慎,现在是不是需要向这位“友商”讨教品牌高端化经验了?

2026-08-17 李宇轩 人工智能 中-英

近日,2021 亚马逊云科技线上黑客松圆满结束。 AI 应用的加速落地,智能驱动的科技创新带来了生活方式的变革与重塑,让美好生活变得有形和鲜活,这场线上举行的黑客松又有哪些值得关注之处呢? 自今年 3 月开启以来,本次大赛以「AI 开发 · 践美好 · 见未来」为主题,围绕「AI + 品质生活」、「AI + 自动驾驶」、「AI + 体育运动」、「AI + 电竞游戏」四大赛题,吸引了多位高校学生、创业者、个人开发者、企业开发者报名,让开发创造无限可能。此外,大赛期间还设置了互动直播,邀请了亚马逊云科技资深技术专家莅临直播间,与开发者们进行了充分的互动交流。 来自全球各大企业和高校的编程爱好者围绕「善用 AI,创造美好生活」、「以不变的 AI 核心,应对万变的自动驾驶场景」、「AI 为体育行业带来的新元素、新思想和新玩法」和「AI 加持,开启电竞加速度」四大赛题,结合至少一项亚马逊云科技产品服务,提交了多个优秀作品。 本次大赛共收到了 12 组选手的提交方案,经历了为期半年的比赛,4 支队伍突破了严格的筛选,获得了决赛入场券。 三强团队 在最后的环节,黑客松的专家评委从赛题完成度、创意度、使用体验、商业价值、展示度、影响力、技术架构多个维度对选手们的作品进行综合考量,评选出了 3 支优胜队伍。 评审团队阵容 7 月 6 日,主办方正式公布了决赛前三名获奖团队,获奖方案分别为「AI 会议便签条」、「减肥辅助机器人」和「多语言交流学习便捷助手」。 获奖方案介绍 在工作和生活中,会议记录的分发和整理是一件费力不讨好的事情,对于想要了解具体开会内容的人,阅读连篇累牍的会议记录很难快速获取自身感兴趣的摘要内容。 冠军团队「AI 会议便签条」利用公开的美国加州圣何塞市 2020 年市政会议记录(共计 45 份会议记录,842 页,1648 个段落,11 万字)作为训练数据集。通常政府的市政会议内容会以 PDF 的格式进行线上公开,人们很难根据具体内容进行检索和查阅。 团队对会议内容进行结构细分和文本预处理,把文本数据转换成便于机器学习算法直接使用的实值向量,并用自然语言处理的算法进行主题模型(Topic Modeling),加入词嵌入(Word2Vec)技术进行分类特征处理,从而让公众可以对市政会议里感兴趣的主题进行跨时间线的检索与订阅,像翻阅便签条一样,无须阅读完整的会议记录而获取感兴趣的细分内容。作品未来的愿景是推广到企业内部,对于各部门会议的具体内容进行片断式的高效分类、检索和订阅,从而节省部分与会者的参与时间,以建立自动化的 AI 系统来提高开会和沟通的效率。 在冠军方案的打造过程中,团队使用了多项亚马逊云服务,最终在一个较短的时间内建立了完善的方案,实现了理想的模型效果。团队特别提到,亚马逊云科技提供了极为全面的云服务解决方案,各种服务器、开发管理和云计算工具,应有尽有。比如 Amazon S3 / Amazon Simple Storage Service / Amazon Elastic File System 可用于储存用户上传文本数据,并添加到原始数据集,重新训练并优化文本搜索算法;亚马逊云科技 Textract 用于提取 PDF 原始数据,对手写文本可以自动识别,并且附带了许多数据分析信息,非常方便;而 IAM / 亚马逊云科技 Key Management Service 可根据团队成员的工种设置不同的应用权限,安全可靠且互不影响,并可共享同一个 Billing 账户,极大地方便了开发管理。此外,AmazonCloudWatch / Amazon Simple Notification Service 用于监控网站异常情况,便于相关人员及时维护,设置好监控指标以后,便可发送邮件或短信自动通知。 本次冠军团队的奖品内容包含:3000 美金亚马逊云科技服务抵扣券、4999RMB 九号(Ninebot)One Z6 单轮平衡车 1 台、3598RMB GoPro HERO9Black 套组 1 套等。 亚军团队「减肥辅助机器人」的产品功能主要有两个:一是信息问询,输入食物或者运动,输出食物对应的热量,以及运动对应会消耗的热量;二是减肥咨询,可通过用户输入身高体重,计算出静息代谢,也可以通过用户输入目标体重和现在的体重,计算出每天应该消耗多少热量,摄入多少热量。 这款产品在打造过程中使用了亚马逊的机器学习云服务 Amazon translate,团队成员介绍说:「这些训练数据都来自于自己减肥时积累的数据。」产品采用了最流行的微信公众号作为产品入口,并在内部分为 NLP 模块、减肥信息结果计算模块与 Amazon translate 几个模块,结构清晰。通过使用亚马逊云科技服务,团队可以方便地直接调用很多成熟的能力,减少开发成本,提升开发速度,对于小型产品的发展非常有用。 亚军团队的奖品内容包含:2000 美金亚马逊云科技服务抵扣券、3598RMB GoPro HERO9Black 套组 1 套、2499RMB 索尼 (SONY) 蓝牙降噪游戏耳机 1 只等。 季军团队「多语言交流学习便捷助手」的设计初衷来源于人们去其他国家旅游或者出差时对于异国语言交流的需求,作品主要实现多种语言的输入及对应的翻译输出。考虑到实际使用时的便捷形式,这款产品除了支持文本输入以外,还支持语音输入。另外,本产品也可作为语言学习工具,让用户能够比较方便地学习和锻炼简单语句的翻译。 在制定方案时,团队使用了 Amazon S3、Amazon Transcribe、Amazon Translate 三项亚马逊云提供的机器学习服务,Amazon Transcribe 负责将存储在 Amazon S3 中的音频文件转换为准确、完整的标点符号文本,不需要开发人员熟悉自然语言处理的相关理论,加快了系统开发的效率,而 Amazon Translate 帮助产品实现了高准确率的文本翻译。产品主要分为前端交互界面和后端逻辑处理部分:交互界面负责接收用户的输入信息,当有输入信息时,触发后端服务进行对应的逻辑处理;后端逻辑处理负责根据输入信息,使用 Amazon 云科技工具进行处理,得到输出结果返回给前端的接收接口。在使用语音输入时,将 Amazon S3 作为数据存取的中转站,设定 bucket 存放和取出语音信息及转换结果。 季军团队的奖品内容包含:1000 美金亚马逊云科技服务抵扣券、2499RMB 索尼(SONY) 蓝牙降噪游戏耳机 1 只等。 此外,所有参赛团队均获得了主办方提供的「阳光普照奖」,奖品包括 300 美金的亚马逊云科技服务抵扣券等。 小到一「码」在手,出行无忧,大到万物互联,智能升级,人工智能已经成为人类生活中的技术底色。在这次圆满完结的 2021 亚马逊云科技线上黑客松中,我们见证了「善用 AI」提升生活品质的应用,也有走向精细化的自动驾驶应用场景探索,以及智能加持的技术创新与文娱、体育场景。未来,秉承着践美好、见未来的初心,亚马逊云科技将鼓励越来越多的开发者参与到赛事之中,用技术创新推动社会的进步。

2026-08-17 李宇轩 人工智能 中-英

近日,法国和美国科学家联合在机器人领域著名期刊Science Robotics上发表了一篇题为《从独立和无意识机器人集群到定向移动灵活超结构机器人》( From collections of independent, mindless robots to flexible, mobile, and directional superstructures )。论文研究表明,该研究团队通过棒状的单体机器人和柔性支架组合成为一种超结构新型机器人,该机器人可以根据不同的地形从而使得自身变形,从而实现越障和运输,可以代替单体机器人实现一些复杂的任务。 单体和超结构机器人的组成 不同于化学中的超结构含义,该研究里提到的超结构指的是基于机器人集群的新型复合结构。椭圆形状的单体机器人的尺寸为 4.5 cm * 1.5 mm,在研究当中,研究人员利用偏转运动产生这种机器人的定向移动。而为了产生偏转运动,单体机器人的软体驱动足设计成了非对称结构,而且嵌入带有电源模块的偏心质量电机以实现驱动。因此,机器人的运动速度通过可以调整驱动频率和电压来进行控制。同时,由于嵌入了光敏粒子,可以通过控制光来控制单体机器人,从而控制超结构机器人的定点移动。 超结构机器人则是由若干个单体机器人和柔性支架组成,通过产生循环振动来产生定向运动。这样一来,超结构机器人的运动能力就取决于单体机器人的输出特性。由于支架的空间灵活性,超结构机器人可以变形并穿过狭窄的空间和实现越障,赋予了该类机器人的空间探索能力。 图 1 | 利用光控制的单体机器人集群到超结构机器人( 来源:Science Robotics) 为了探索其运动性能,研究人员对这种超结构机器人在对复杂几何形状空间的探索进行了研究,分别完成负载、越障、清障等任务。在实验中,为了追踪这些单体机器人的运动轨迹和获取相关运动数据,研究人员把单体机器人涂成了黑色,便于观察。 超结构机器人在直通道中的运动 在多个单体机器人开始偏心旋转运动具体在一起,在一定时间后,这种偏转运动转化为超结构机器人才开始运动,实现简单的直线运动: 动图1 | 超结构机器人在直通道中的直线运动( 来源:Science Robotics) 而且在椭圆型通道中,超结构机器人可以适应地形以进行拐弯: 动图 2 | 超结构机器人在直通道可以适应地形运动( 来源:Science Robotics) 图 2 | 超结构机器人在直通道中的运动能力( 来源:Science Robotics) 超结构机器人穿越不同宽度通道 实际情况下的地形是十分复杂的,越障能力的提高对于进行空间探索机器人的发展至关重要。因此,研究人员研究了超结构机器人在越障时的能力,而能力评价标准之一是越障时所花的时间。 动图3 | 超结构机器人穿越狭隘地形( 来源:Science Robotics) 此外,研究人员对比了实验和仿真下的超结构机器人在不同宽度通道下的运动能力: 图 3 | 仿真和实验环境下的超结构机器人越障能力对比( 来源:Science Robotics) 不仅可以实现自身的运动,超结构机器人还可以通过单体机器人的运动来实现负载( carrying load ),而且这种超结构机器人的负载能力可以达到其自身重量的一半: 动图 4 | 超结构机器人负载进行运动( 来源:Science Robotics ) 而且还可以负载的同时,由于自身结构的优越性,轻松实现越障: 动图 5 | 超结构机器人负载同时实现越障( 来源:Science Robotics) 图 5 | 超结构机器人负载穿过圆柱体(障碍)( 来源:Science Robotics) 单体机器人和超结构机器人的应用 当路面上存在障碍物(空心圆柱体)时,机器人还可以实现对路面进行清理。研究人员对比了两种类型的机器人在清障效果。 动图 6 | 单体机器人清障( 来源:Science Robotics) 可以发现单体机器人的清障能力有限,而装载了 20 个单体机器人超结构机器人在不到 40 秒的时间里把空圆柱推出场地,实现高效率的路面清障: 动图 7 | 多体机器人对路面进行高效率清障( 来源:Science Robotics) 可实现光控驱动的超结构机器人 当没有光线时,超结构机器人处于静止状态。当灯打开 10 秒后,由于单体机器人嵌入了光敏粒子,光照射到光电晶体管上时,机器人的电机被驱动,从而驱动组会产生顺时针方向的偏转运动,机器人之间发生碰撞,形成一个移动集群,最终实现超结构机器人的驱动: 动图 8 | 光线控制机器人进行运动( 来源:Science Robotics) 同样,光线控制下的超结构机器人也可以实现越障和用于清障: 图 6 | 光线控制下的超结构机器人用于清洁场地和越障( 来源:Science Robotics) 该团队对超结构机器人进行了研究,包括对机器人进行负载穿过障碍或清理场地。还提出了使用光对超结构机器人进行基本控制。发现可以通过使用结构简单,价格便宜的单体机器人能够制造出具有空间探索特性的灵活的超结构机器人,为未来机器人清理障碍物,通过收缩通道等难题提供了一定的研究价值。 -End- 赞助本站

2026-08-16 李宇轩 人工智能 中-英

人工智能(AI)在医疗行业的应用一直都受到广泛关注。AI研究如何在遵循临床研究范式的基础上寻求变革创新?如何看待AI在医疗实践中的伦理问题?谁将占据AI医疗的制高点? 针对上述问题,8月29日,《新英格兰医学杂志》(NEJM集团)与嘉会医学研究和教育集团在上海举行了一场医学AI研讨会,来自全球顶尖的医学专家和业内领先企业分享了AI在医学应用中的实际应用案例以及未来面临的挑战。 AI系统的能力如何评估? NEJM集团编辑德莱森(Jeffrey Drazen)教授从AI医疗研究出版论文入手,他援引一篇发表在2020年4月出版的《新英格兰医学杂志》上的论文。该研究使用人工智能检测视神经乳头水肿的效果,通过计算接收工作特性曲线下面积(AUC)、灵敏度和特异性来评估视盘外观分类的性能,并与神经眼科医生的临床诊断参考标准进行比较,AI在某些方面的工作甚至超过了眼科医生。 德莱森教授表示,机器学习应该对大量眼底照片进一步研究,这些照片的资源非常宝贵,一方面是因为通常受到隐私保护的限制,另一方面是眼科器械设备的高昂价格限制了资源的普及性。 复旦大学上海医学院眼科学与视觉科学系主任、中国研究型医院学会眼科学与视觉科学专委会主委孙兴怀教授对第一财经记者表示:“医学成像中人工智能算法正在实现爆炸式增长,人工智能在眼底病筛查上具有很大潜力。但AI要在临床上应用好,一定要有一个非常强大的神经网络计算机团队的支持。” 他同时指出,人工智能靠大量的前期病例图片输入进行识别面临的局限性。“因为临床上疾病是千变万化的,以前没有输入过的,它是识别不了的。”孙兴怀表示,“有些疾病不只是表现在眼底,还要结合眼部其他的改变来综合分析判断。人工智能只能给出一个倾向性判断,最终还是要有经验的医生综合分析来决定。” 医疗涉及到人类的健康,对AI医疗的系统进行评估的重要性不言而喻。德莱森教授表示,未来开发出一套具有针对性的关于AI介入诊断精确性的评估指南很有必要。 “在评估AI系统能力的时候,一方面是评估AI的知识和能力,需要满足临床设计的目标;另一方面,AI系统应该是一个可以不断进化和升级迭代的系统,让技术的发展适应人类。”德莱森说道。 麦考瑞大学教授、澳大利亚健康创新研究所所长Enrico Coiera表示,AI系统大致可以分三类,包括辅助医生的AI系统,比如读心电图;另一种是可以替代医生的AI系统,比如筛查眼底照片;第三类是AI系统将来可以做的医生目前还做不了的事情。 “我们认为评估AI的性能,也应该将人机互动的性能结合起来进行评估,而不仅仅是评估算法,AI系统如何最大程度上帮助人类医生,这是一个更为复杂的问题。”Coiera教授表示。 他还强调,医疗AI应该不仅仅作为一种通用的技术,而是可以作为专业行业的定制化技术,这对于专科而言非常重要。 医疗AI如何获得信任? 专家普遍认为,AI在医疗领域的广泛应用的前提是要在工业界、科学界、医生以及患者群体中获得信任。有趣的是,多位专家将AI医疗系统和无人驾驶系统进行了类比,无人驾驶上路要先考“驾照”,那么AI医生要上岗也需要“持证”。此外,万一AI系统发生事故,那么相关的责任认定也需要有所规范。 “就像自动驾驶一样,我认为现在的AI还没有完全成熟。”中科院院士、复旦大学附属中山医院心内科主任葛均波教授表示,“作为临床医生,在手术过程中,我们希望AI能够指导医生进行非常精确的操作。目前看来,这一功能还要通过机器深度学习,提高应对各种复杂状况的应变能力。” 今年3月,葛均波在西门子医疗的手术机器人系统途灵的辅助下,完成了中国首例机器辅助冠脉介入手术。在机器辅助手术中,葛均波通过手术室外的操纵杆,隔室操控机器人进行手术,让导丝迅速通过复杂的病变,微调冠脉支架系统的位置并释放。 葛均波认为,机器人目前的操作还无法达到业内专家的水平。不过他仍然相信,AI系统一定能在慢病随访、疾病筛查、基层医生诊断等领域发挥作用,并且随着未来在感知方面能力等提升,可以最终达到人类顶尖医生的水平。“AI现在达不到预期,并不意味着我们就要停止对它的研究了。”葛均波表示,“未来AI一定会在精准性和标准化方面发挥重要作用。” 西湖大学特聘研究员郭天南表示:“对AI的信任是基于AI所创造的价值,但人们也不应该对AI有过高的期望值,仍应循序渐进,同时加大AI在临床的渗透率。” 专家还表示,目前很多基于深度学习的系统仍是一个黑箱,涉及到很多参数,在相关领域数据共享方面的工作进展仍然非常有限。 AI实验室AI医疗首席科学家姚建华表示:“虽然从公司盈利的角度来看,不太可能把所有的源代码都开源,但是用于AI系统训练的数据是很重要的,我们使用什么样的数据来进行模型的训练和检测,数据是否有偏见,是否符合临床实践的价值,这些应该是适合发布给公众了解的。” 医渡云首席人工智能科学家闫峻表示:“保证数据的质量非常重要,我们需要的是用于临床研究的结构化和标准化的模型。现在即便是非常好的医疗机构提供的数据,很多也不能直接拿来面向训练,需要自然语言处理技术来识别,识别的过程中需要很多医学逻辑。” 科大讯飞一位医疗AI负责人强调了医疗AI系统的可解释性。“AI医疗的智能化系统如果提供了服务信息,那么医生需要获得解释;如果AI系统作为一个辅助诊疗系统,那么本身也需要可解释性。”这位负责人表示,“在样本的可解释性技术路线之上,我们再对数据做深度分析,寻找输入和结果是否强相关的,最后还需要结构化的医学诊疗知识。” 赞助本站

2026-08-16 李宇轩 人工智能 中-英

“Linux是否抄袭Unix”之争 博雯 发自 凹非寺 量子位 报道 | 公众号 QbitAI 这场状告Linux抄袭Unix的官司,终于在20年后结束了。 起诉方为SCO(Santa Cruz Operation)公司,主要业务为运营并销售UNIX及其相关产品。 而在昨天下午,代表SCO公司的TSG集团与IBM达成了和解: SCO将放弃,并再也不会对Linux进行违反Unix或Unixware知识产权的指控。 同时,IBM将也支付1425万美金(折合人民币9217万元),作为对SCO的全部赔偿。 “谁背叛了联盟?” 到底什么官司,居然打了20年? 故事开始于1998成立的一个联盟,Project Monterey。 这是一个由IBM、SCO以及其他公司所创立的一个项目,目的是开发一个适用于多种硬件平台的UNIX版本。 也是Linux正在做的事。 到2001年,IBM的Big Blue超级电脑已经创建了一个类Unix的AIX操作系统(实验版本)。 而这一系统使用了一些SCO代码。 这时候,IBM认为Linux才是未来,于是退出了Project Monterey项目。 但SCO表示了反对,理由是IBM属于联盟,或者说属于SCO的知识产权贡献给了Linux。 于是,2003年3月6日,SCO公司一纸诉状将IBM告上法庭: SCO对Unix和UnixWare操作系统源代码具有所有权,而Linux 2.4.x和2.5.x是 Unix的未经授权的衍生物,或者说是“抄袭”行为。 也就是说,IBM传播Linux代码这一行为造成了严重侵权。 同时,他们还致函全球500强企业,警告他们如果继续使用Linux,将可能承担法律责任。 并且还表示: 未来可能拒绝通过Linux使用者的许可证申请。 IBM和Linux分销商Red Hat也不甘示弱,反手一个诉讼将SCO与相关公司告上了法庭。 于是,这场围绕知识产权的战争就此打响。 起诉方几经变换,最终千万美金和解 一开始,SCO就这一行为向IBM索赔10亿美元,而这个数字后来增加到了50亿。 在旷日持久的激烈争斗中,甚至一度有“SCO对IBM的诉讼可能终结Linux”的说法。 连起诉方几经易主,都没能阻挡这场上诉、结案、推翻、再上诉的重复战争。 没错,最初发起上诉的SCO公司,其实在2007年就已经申请破产。 在2011年时,SCO将资产出售给了Xinuos公司。 Xinuos的CEO曾在2016年表示: 我们是购买产品的投资者,并没有买到对IBM进行诉讼的能力,我们对此完全没有兴趣。 但在5年后,Xinuos又一纸诉状,控告IBM和Red Hat公司版权侵权和反垄断: 而今天站在起诉方终结这场官司的,是代表SCO公司的债务人:TSG集团。 最终,IBM以缩水的赔偿金,换来了Linux未来再受“违反知识产权”指控的可能。 至于为什么最终达成了和解? SCO公司一方的法律代表Stanley B. Tarr表示: 想要拿到索赔,就必须向评审团证明多年前发生的的事件构成了不正当竞争,并造成了SCO的损害。 并且,即使上述的申诉成功,SCO所要求赔偿的损失金额也并不确定。 事实上,受相关损失限制条款、以及IBM反诉的影响,陪审团不可能作出对SCO有利的判决。 但不论背后的原因究竟为何,至少在20年后,SCO vs. IBM的争斗终于走到了终点。 参考链接: [1]https://www.zdnet.com/article/after-almost-20-years-the-sco-vs-ibm-lawsuit-may-finally-be-ending/ [2]https://en.wikipedia.org/wiki/SCO_Group,_Inc._v._International_Business_Machines_Corp. [3]https://www.theregister.com/2021/08/30/sco_tsg_vs_ibm_settlement/

2026-08-16 李宇轩 人工智能 中-英

模拟人眼,所见即所得 杨净 发自 凹非寺 量子位 报道 | 公众号 QbitAI 最近,大家都在讨论元宇宙。那是否有想过,在元宇宙里应该会如何购物? 是像这样,在手机端就能看到跟线下店一模一样的完整实物? 还是足不出户就可以美美的线下店? 又或者是,看上哪款就试哪款? 旁友们,是不是以为这就是搞个3D模拟,或者电影特效那种后期制作? NoNoNo,其实都不是。 这个用到的是光场技术,模拟人眼,所见即所得。 这技术可来头不小。 前段时间谷歌I/O大会亮相的裸眼3D通话,就是基于这个原理—— 将空间里人眼看到的光线进行采集重建,重现真实的三维世界。 如果把它用到我们日常的穿衣住行上,简直“高质量”了有木有! 事实上,已经有公司做到了。 而且还是一家中国公司——奥本未来,哈佛团队班底,那几张画面正是他们与几大品牌商的合作。 那么具体是如何实现这种逼真感呢? 什么是光场技术? 提到光场,就要说到我们的视觉感知。 人类能看到所在的世界,就是通过人眼,来接收现实场景中的光信息,从而进行感知。 如果我们能把空间中的光都记录下来,那我们就可以从任意角度、任意路径,去观察,去体验这个空间,就像到现场一样。 元宇宙电影《Ready Player One》里,主角进入前人的记忆信息,各种漫游,甚至查看细节来破案,要真的实现这样的场景,就需要用到光场技术! 光场技术,记录了七个维度的信息。 他们分别是场景中任意点的位置(x, y, z),用来表示它发光方向的水平角度和垂直角度(θ, Φ),发出光线的波长(λ)以及时间(t)。 别小看这七个维度,它就已经能对场景100%光线描述。 人看到什么,它就捕获到了什么,所以能根据观看需求任意调整视点、焦距和景深。 实际上光场技术之外,也有其他的方式来记录空间信息。 比如三维建模。相信玩3D游戏的旁友很熟悉了。要实现逼真的3D模型,且不说逼真的实时光照需要巨大的算力,单就光照、材质、几何等要有极高的保真度,就已经极具挑战性了。 或者使用一些摄影测量或者三维扫描的设备,但更多细节无法捕捉到,离电影里逼真的渲染效果相去甚远。 所以光场采集无疑是短时间内最快实现商业化的解决方案。 奥本未来就是研发的这项技术,有这几方面的的优点。 首先,采集设备比较成熟。只要一部手机就可以做光场的采集。 以往学界做光场采集,动辄就是几百个相机的阵列。 要么是由几百个小摄像头集打造的光场相机。 但奥本未来的技术团队表示, 采集过程可以使用任意相机,苹果手机后摄也可以。 但为了保证照片质量,一般使用单反相机。 此外,技术人员还透露,通常一件典型的商品只需要采集一两分钟的视频或者一百多张照片。 以往光场数据十分庞大,想要在屏幕上完整渲染出来,需要高性能主机才能做到。但如今通过压缩算法,普通网页也能实现逼真效果。 比如,京东APP上试穿三维运动鞋,就用了奥本未来的光场技术。 还有在unity做的三维游戏也用到了光场模型,动起来的那种。 正是基于以上这些技术突破,奥本未来让真实空间和物体的整套流程变得更加高效。 目前,奥本未来自研了一款光场捕捉器,最小化人为干预,只需连接普通相机,即可快速自动化完成对常见物体的光场采集。 在面对海量SKU的电商行业,从采集、重建、部署到上线平台整套下来只需要一天就可以完成。 技术人员透露,当中成本只包括两部分:拍摄照片的人力和云端重建计算的算力,相当于传统三维建模的十分之一。 不过,他们坦言,当前仍然具有诸多挑战,比如,大规模场景重现。 跟高奢品牌打交道的奥本未来 最后,再来简单介绍一下奥本未来。 据官方介绍,它是一家用前沿技术来解决三维内容供给,并致力于将实时三维交互推向更多行业背景的技术公司。 团队方面,奥本未来由哈佛大学海归团队于2016年创立,创始人兼CEO雷宇博士是一名连续创业者,博士毕业于哈佛大学应用物理及材料科学系。 CTO沈方阳博士毕业于北航,虚拟现实领域博士,曾在微软亚洲研究院、哈佛可视化计算组(Visual Computing Group)从事计算机图形学和视觉方向研究 此外,技术团队还来自于哈佛大学图形学实验室、北航虚拟现实国家重点实验室、微软亚洲研究院、清华、北大、北邮等各高校机构人员组成。目前也在积极广纳英才,比如图形学程序员了解一下? 融资上,团队们上轮完成了Pre-A轮融资,由高通、京东领投。此前的天使轮还包括华创资本以及多家清华系基金等投资机构。 截至目前,奥本未来已经落地服务超过30家品牌客户,都是大家耳熟能详的著名品牌。 当然也有一些接地气的平台合作伙伴,比如京东、微信小程序、天猫等主流电商平台。 而之所以跟品牌方、电商平台打交道,源于2019年的洞见。 随着用户对电商购物的体验要求越来越高,3D互动导购内容和场景开始大量增长, 相比与以往图片、视频等有限视角,3D素材、3D实时交互以及AR试穿试戴素材相关的需求只会越来越高。 直至现在,电商随处可见3D内容场景,比如商详页、频道页、店铺页、直播间、搜索页、会场页等等。 对于品牌电商来说,一整套低成本、门槛低、轻量级的三维内容方案成为他们的首选。 比如,便携式采集技术和低门槛的三维编辑工具可以帮助品牌与内容制作者节省制作的成本; 整个编辑存储发布都在云端,有助于多方协同;整个方案可以支持多平台发布的属性。 这么说起来,奥本未来所做的更像是一套底层基础设施,推动整个购物体验数字化、沉浸化、智能化。 如今元宇宙概念的爆发,似乎也印证了奥本未来之前的选择。 至于实现真正的元宇宙,需要更多像奥本未来的公司来共同打造技术底座。

2026-08-15 李宇轩 人工智能 中-英

Dataset Biases: Institutionalized Discrimination or Adequate Transparency? A review of the efforts performed by the US Mortgage Disclosure Act […] Why would we need such a law? Prior to Congress’s enacting HMDA in 1975, the public raised considerable concerns about mortgages — or, more importantly, the lack thereof — in some urban, often minority, neighborhoods. Certain areas seemed to decline, in part because their residents were not able to obtain home mortgages. (ClevelandFed) This was one of the sad realities of certain American population centers in the 1970s. Access to capital remained difficult, and social mobility for at-risk neighborhoods was quasi-non-existent. This difficulty was accentuated by some believed to be institutionalized racism in the banking system. “Congress believed that some financial institutions had contributed to the decline of some geographic areas by their failure to provide adequate home financing to qualified applicants on reasonable terms and conditions.” (Wikipedia) Therefore, a motion was brought forward to Congress to support transparency across all lending practices. They passed the Home Mortgage Disclosure Act of 1975 for mandatory reporting of all loan applications, and then the Community Reinvestment Act of 1977 for encouraging financial institutions to help meet the credit needs of their local communities. There is a very clear mandate to the HMDA, as explained by Investopedia: In general the primary purposes of the Home Mortgage Disclosure Act and Regulation C are to monitor the geographic targets of mortgage lenders, provide an identification mechanism for any predatory lending practices and to provide reporting statistics on the mortgage market to the government. The HMDA helps to support the community investment initiatives sponsored by government programs, with HMDA contributing to the oversight of the initiatives through statistical reporting. HMDA also helps government officials to identify any predatory lending practices which may be affecting mortgage loan issuance. HMDA submissions also provide a means for analyzing government resource allocations and ensuring that resources are appropriately allocated to fund community initiatives. Therefore, financial institutions must report on their lending practices specifically by reporting not only all of the loans issued but also every loan application with the associated metadata of the applicant(s), such as race, gender, and neighborhood, as well as if the loan was approved or denied. From a regulatory perspective, there was now a lens through which violations to social equality can be tracked and penalties applied. This also means that there is an intentional bias in the data — that of institutionalized racism across the banking sector, and all of its manifestations. Exploring the Data Let’s explore the Home Mortgage Disclosure Data Files (1981–2014) from the US Archives to get a sense of what has been reported, and why. (If you want to explore this dataset at home, you’ll also need a few more datasets, such as the relevant census data. Luckily, the superstar team of librarians at the US Archives packaged everything up for us in a single page. Thank your local librarian.) 1981–1990 By exploring the data, we see that the 1981–1990 period was primarily focused on tracking veterans, many of them returning from Vietnam. A lot more of the questions and data pertains to VA applicants, and if they had requested financial support to apply for the loan. The encoded format is quite straightforward, just needs a bit of mapping: Overall, lots of VA considerations. 1990–2014 A big change was introduced in 1989 to the tracking priorities. From the FFIEC: “[…] In 1989, the Federal Reserve Board revised Regulation C, to incorporate amendments contained in the Financial Institutions Reform, Recovery and Enforcement Act (FIRREA). The FIRREA amendments accomplished the following: expanded the coverage of HMDA to include mortgage lenders not affiliated with depository institutions or holding companies; required reporting of data regarding the disposition of applications for mortgage and home improvement loans in addition to data regarding loan originations and purchases; and required most lenders to identify the race, sex, and income of loan applicants and borrowers.” As expected, the file mappings changes significantly, bt now follows a pattern usable until 2014: This means that we can now tell, for every application: The race and gender of the applicant; the race and gender of the co-applicant; the neighborhood that they were in; the purpose of the loan (house or repair); the list of reasons why a loan might’ve been denied; the total number of forms required to be submitted; and each and every one of the dates of the application process. Was a loan application intentionally slowplayed? Was a home repair denied that would’ve otherwise been approved? What about same-sex co-applicants? All visible. (Note: we haven’t made our processed data available yet, as we’re still processing and investigating it. There are many reporting errors, like state acronyms instead of state codes, erroneous census tracks, and many other issues that still require a tremendous amount of preprocessing of the data to make sense of decades-long shifts in social norms. It would be unwise to come to any conclusions with it at this time. ) Over the years, there’s more and more adherence to the program across most banks. The digitization wave that spread across the US can be verified year over year. There’s also a distinct shift downwards of less and less mortgage loan applications, stemming most likely from the subprime mortgage crisis: Great so far, isn’t it? Now, here’s the double-edged sword about this data: The same data that can be used to protect citizens against predatory bank practices can be used as a discriminatory weapon against them if data scientists do not understand the data they are using. How so? This dataset now has within it every bias, preferential treatment, and erroneously or maliciously declined loan application. Every possible discriminatory event is permanently recorded. There is no sanitizing possible of this data. The HMDA dataset is biased — absolutely. Should this data even exist? By any measure, this data set can be considered invasive and discriminatory. Canada, usually considered a (slightly) more tolerant country, has a different approach to tracking race and racism. As explained in The Conversation: Canada’s anti-racism strategy, which draws on decades’ worth of research, states that race is a social construct. There is no basis for classifying people according to race, but racial bias and discrimination have very real effects. The question is: How do we get relevant data from the census and other surveys on the impact of systemic racism? Statistics Canada tries to gather this information without directly asking about race. Race-based data is needed, says Jean-Pierre Corbeil, a diversity specialist at Statistics Canada. But he wonders whether that actually requires referring to race on the census. Historically, the government has been reluctant to ask directly about race, which has led to a lack of disaggregated data. After the Second World War, the census used indirect methods of estimating the non-white, non-Indigenous population through racial proxies like language or ethnocultural origin. (Note: one of the major complaints against Canadian multiculturalism as a pillar of civil society is that it allows for people to classify others by ethnicity or nationality, under the veil of a self-assigned permission structure, but that’s a topic for later.) So, this means some countries refuse to track the race and gender of bank loan applicants, as measuring racism fundamentally amplifies its very divisive nature of bucketing people into categories. And so, looking back at the HMDA data, the question that can be asked is, “should this data even exist?” I believe the answer is a strong yes. The Act has yielded tremendous protection to the public, especially with a major investigation back in 2005. From the Buffalo News: New York Attorney General Eliot Spitzer has fired his latest salvo, launching a preliminary probe into mortgage lending practices at eight major banks in the state, including HSBC Bank USA. […] Investigators are trying to determine how the banks price their loans, and if fees and interest rates are being applied fairly, or whether there’s racial discrimination. […] Citigroup already agreed in 2002 to pay $215 million to settle allegations by the Federal Trade Commission that Associates First Capital Corp. — which Citi bought in 2000 — had engaged in predatory lending. Associates’ rival Household International — acquired by HSBC Holdings Plc in 2003 — paid $484 million in 2002 to settle similar charges by all 50 states in the largest consumer settlement ever. However, although this data has yielded success, it’s up to the data scientists to understand what are the right questions to ask, and more importantly what not to ask. For instance, trying to build a mortgage loan approval pipeline from this data is a horrendous and terrible idea — you now have a racist, homophobic bot. However, if you wish to investigate the very nature of institutionalized discrimination by investigating outliers, then you have a tremendous weapon at your disposal. How you wield that weapon is up to you. Additional Reading Revisiting the CRA: Perspectives on the Future of the Community Reinvestment Act, A Joint Publication of the Federal Reserve Banks of Boston and San Francisco, February 2009. Privacy Impact Assessment of the Home Mortgage Disclosure Act Data Repository System, Board of Governors of the US Federal Reserve, November 2020. I wrote the law Bloomberg blames for the financial crisis. He’s wrong., Robert Kuttner, op-ed in the Washington Post, February 15, 2020. History of the HMDA, Federal Financial Institutions Examination Council, January 2018. Happy Consulting! -Matt. If you have additional questions about this article or our AI consulting framework, feel free to reach out by LinkedIn or by email. Other articles you may enjoy How Does AI Create Value? Implementing a Corporate AI Strategy Outlier-Aware Clustering: Beyond K-Means Rorschach Tests for Deep Learning Image Classifiers

2026-08-15 李宇轩 人工智能 英-中

By John P Desmond, AI Trend Editor Experiences with AI and machine learning at CVS Health and St. Luke’s Health System in Boise, Idaho, are having practical benefits to the two organizations. CVS Health is learning how to scale AI applications using machine learning, especially through the house of machine learning operations (MLOps) tools, according to Nels Lindahl, director of Clinical Decision Systems, speaking in a virtual session at the recent Ai4 Conference held virtually recently. And St. Luke’s Health Center put a COVID-19 prediction program, a supply chain purchase engine and a demand-based staffing application into initial production using AI and machine learning, said Dr. Justin Smith, senior director of advanced analytics at St. Luke’s, also at a recent Ai4 virtual conference session. “We are at an MLOps tipping point, where ML has a growing production footprint, with adoption picking up pace and awareness and understanding at an all-time high,” stated Lindahl. “ML tech can now deliver; people are seeing real use cases in the wild and having them grow; it’s real.” The three primary “ecosystems” for building out an AI footprint are Amazon Web Services (AWS), Microsoft Azure and Google Cloud Platform (GCP), he said. Developers are using open source tools and some they develop on their own to deliver value to their organizations. He highly recommended having an ML strategy, with executive sponsor, defined budget resources, and a repeatable process to ensure the same approach can be replicated throughout an enterprise. “You need a narrative that helps you push things forward. When you can deliver on a use case consistently, you want to be ready to go,” he said. In his professional role today, Lindahl is directing IT application development. He has delivered several complex process automation projects within the clinical decision space of pharmacy benefit management. Tracking Popular API Services on GitHub In April, he was tracking 19 external API services on axes of scale versus maturity, with for example AWS Enterprise Search and GCP Vision API rated high on the scale and maturity axes, so in the upper right quadrant. He recently updated his chart and is now tracking 49 external services. “The number of APIs out there is growing every day,” he stated. The external services are becoming more specific, so that instead of one large vision API, it is breaking down into more specializations. “The exact thing you want to do is probably closer to something you can put into production from your API, ” he said. “The AI is out there and ready to go.” He emphasized again the need for a development organization to have an overall ML strategy. “Just because you can go out and get an API, does not mean that is the right thing to do for your organization. It might be cool technology that is amazing, but there might be no return on that investment for your customer,” he said, adding, “Part of your ML strategy must be about purpose, replication, and reuse. Those are going to be at the heart of getting value back for the organization.” He recommended looking at the GitHub Trending page, to see the popularity of different services. He tracks them over time, using indicators such as “stars,” which means people who have indicated a continuing interest, like a bookmark. In April, for example, TensorFlow registered over 154,000 stars and Pytorch was at 47,000; those attracted the most interest. Lindahl is also tracking MLOps: machine learning operations tools that support practices to maintain ML models in production reliably. His tracking of MLOps on GitHub shows 10,000 stars for Kubeflow, a machine learning toolkit for Kubernetes, which is an open source container orchestration platform. What surprised Lindahl more was the rate of growth for MLReef, an open source MLOps platform with a focus on collaboration. “The number of people who downloaded it and are actively using it is going up really fast,” he said. He is researching it to see if he can understand what is driving the increased use. Three ML Projects at St. Luke’s Delivering Business Value in New Ways The emphasis of Dr. Smith at St. Luke’s was more on current delivered projects, and less on their technical underpinnings. The St. Luke’s complex has eight medical centers serving about one million patients. Dr. Smith described three projects. The first was an application to predict the rate of spread of COVID-19, so that the hospital could do some better planning. “We were receiving all kinds of wild forecasts for what to expect,” he said. “We wanted to know if we could predict how many patients we would have in our ICU and on general hospital floors.” The difficulty was that the forecasts were so different. Idaho had not experienced a wave of infections that happened in New York, for example. The team decided to document what they did know, then produce from that data two sets of forecasts: a short-term forecast for one or two weeks, and a medium-term forecast of up to 30 days. They used a new technique, XGBoost, an open source software library popular in applied machine learning for structuring tabular data. It implements gradient-boosted decision trees, which builds a regression tree in a stepwise fashion, measuring the error of each step and correcting it in the next. “It got us to within five or six patients a month out,” Dr Smith said, using variables including the inpatient census and the positivity rate from their own health system. “We controlled 90% of the testing, so we had good data on positivity rates,” Dr. Smith said. “We showed with strong accuracy whether we would be increasing or decreasing the patient census.” For the supply chain, the team was asked if they could develop a “purchasing engine” that could achieve savings with more optimal purchasing from the 100 vendors and 400,000 products in the supply chain. Certain volumes of purchase qualify for better pricing, posing an optimization problem. “With such levels of complexity, it’s too large for a human. You can’t solve it on Excel,” Dr. Smith said. Breaking Analyst Habit of Sticking with Familiar Vendors An examination of the practices of hospital analysts involved in purchasing, showed that many operated in familiar territory, often choosing vendors they knew and had dealt with in the past. Using advanced analytics, the team wrote algorithms that generated “billions” of scenarios. Providing a view of that data was challenging. “We don’t want to show everything or too little either,” Dr. Smith said, The team took the approach of eliminating the non-viable scenarios, sticking to spending within existing agreements, and to spend the required volume with each vendor to get the best price. The team also tried to minimize the number of transitions, in which a product would be purchased from a different vendor. The analysis allowed the team to select opportunities representing the lowest number of transitions and the highest potential for savings across the entire health system. “It’s very powerful,” Dr. Smith said. “It’s being rolled out across the enterprise.” The new system will require renegotiated contracts with a number of suppliers, so Dr. Smith expects it will take several years to become fully implemented. The third project was around demand-based staffing. “We don’t want to be overstaffed and we don’t want to be understaffed,” he said. “We want to optimize our staffing to match historical patient demands.” While COVID did “wild things” to the patient census in 2020, that data was included. The system tried to avoid high labor costs associated with “on-call” roles, and also having to send medical professionals home when demand is low. “The solution we created, which is still rolling out, is based on algorithms that offer an optimized schedule that matches staff and census according to predicted demand based on the last three years of data, including 2020,” he said. The system was focused on nurses and their support staff to start. Using a visual picture of the forecasts via Power BI, the business analytics service from Microsoft, human professionals were able to adjust the recommended staffing based on “front-line knowledge.” That would be, for example, knowing there is a football game, a music festival and a rodeo all in the same weekend, something the machine learning algorithm might not pick up. “The output might be that we need six RNs to cover the 7 am to 7 pm shift, but just four for the 9 am to 5 pm shift,” Dr. Smith said. The system is beginning to be deployed and can project weekly staffing levels needed for four to six weeks out, he said. “It’s very dynamic, and senior executives can look and see how we are doing in hospital staff,” Dr. Smith said. Learn more at the recent Ai4 Conference, at the GitHub Trending page, at Nels Lindahl and at Dr. Justin Smith.

2026-08-15 李宇轩 人工智能 英-中

盛夏百草园,药垄纵横。7月17日,南京中医药大学“脉络寻源实践团”走进江苏食品药品职业技术学院百草园种植基地。从“看图识药”到“田间辨药”,这群青年学子聚焦淮安盱眙道地药材——野马追,践行“从田头到床头,构建振兴国药全产业链思政育人共同体”的理念,试图为这株江淮本草留下真实可感的记录。 识药于野:俯身草木间,辨一叶一茎 步入百草园,科普展板最先留住了脚步。淮山药、前胡、草决明、白术……图文交错间,药材的形态与习性渐次铺开。队员们边看边记,提前圈出“撞脸”特征。纸上有了底,下田便不慌了。 跨过田埂,纸上的名字终于有了真实的触感。薏苡茎秆挺拔,草决明叶片舒展,乌蔹莓沿竹攀援,西洋参隐于林下绿荫。遇到两株藤本难以区分,队员们便凑近反复比对:有人拍下叶脉特写,有人蹲在垄边勾画轮廓。一次次的弯腰触摸,让那些原本停留在图片里的冰冷药名,终于在指尖长出了温热的根茎与枝叶。 问道于土:一味野马追,藏着淮地密码 行至一片菊科植物前,老师停下脚步:“这就是野马追,淮安盱眙的道地药材。”队员们顿时围拢。此前在展板上见过图片,可实物摆在眼前,仍有人误以为是加拿大一枝黄花。正是这个“误会”,撬开了追问:二者到底怎么分?“野马”之名有何来历? 老师蹲下身,从叶片锯齿、茎秆绒毛一路讲到花期差异,又沿着种植周期、田间管护、采收加工,把药材的一生徐徐展开。 话题渐渐从植株本身,延伸到了“道地”二字的分量。盱眙被《中国药典》认定为优质产区,并非偶然。淮河冲积土壤、温润水热与世代农法,共同锁定了有效成分。这恰恰印证了“田头”种植的严苛标准,正是为了守护百姓“床头”用药的安全底线。“道地”远非地理标签,而是自然与人文的长期协同。认识一味本草,不能止步于“有什么用”,更要追问“从何处来、为何扎根”。 建档于案:让乡土记忆走出田野 踏访结束,队员们围坐整理笔记。“一份本草档案,该记录什么?”原植物辨识、产地环境、栽培加工、产业现状……大家逐项梳理,田间照片同步归档。 座谈中得知,当地虽已形成产加销链条,但种质资源保护与公众科普传播仍有提升空间。队员们意识到:建档不是目的,传播才是归宿。让书页外的草木走进孩童视野,让水土故事可感可传,才是档案的灵魂。下一步,团队将结合此次踏访所得,把植株形态、辨识方法转化为科普内容,待这些一手材料走进暑托班与社区课堂,孩子们看到的将不再是枯燥的药名,而是真实生长于江淮大地的本草。 归途|笔墨落处,草木长青 离开百草园时,风从药垄间穿过,叶片轻轻翻动。队员们从展板认药,到田间辨药,再到案头梳理,也愈加明白:本草档案记录的不只是名称与功用,更是一株植物所依存的水土与凝结的经验。 丙午马年,“野马追”的名字恰有一份追风向前的意趣。一株本草能走多远?答案不在典籍里,而在田垄间,在每一次俯身辨认的专注里,在每一份被认真记录的生长数据里,也在未来孩子因一张科普卡片而记住它的瞬间里。随着这份档案的初步落成,一株淮上本草在青年笔端获得了新的生命。这也正是“振兴国药全产业链思政育人共同体”中,青年学子应有的担当——用脚步丈量土地,用笔墨传承国粹,让国药瑰宝在新时代焕发生机。 文字:李忠泽 图片:郑宇恒 免责声明:市场有风险,选择需谨慎!此文转自网络内容仅供参考,不作买卖依据。

2026-08-14 张明 教育资讯 中-英

The Quora Question Pair Similarity Problem A beginner’s journey through the various life cycles of a problem on Kaggle. This is my first case study, so you can expect a beginner-friendly data analysis and model building. I have used only the classical machine learning models for this problem. However, working on this case study was a great learning experience for me. And in this blog, I will try to share with you as much as possible. In the blog, I will write only the summary. You can view the full notebook here and you can view the code on github. And to all the experienced folks out there, I would love your feedback for future case studies. ?? Table of contents Introduction Business Objectives and Constraints Data Overview Business Metrics Basic EDA Data Cleaning Feature Extraction EDA with Features Featurization with SentenceBERT i. EDA on new features related to SentenceBERT Data Pre-processing Training Models i. Support Vector Classifier ii. Random Forest iii. XGBoost iv. Another XGBoost ? Final Thoughts References Introduction Quora is a platform for Q&A, just like StackOverflow. But quora is more of a general-purpose Q&A platform that means there is not much code like in StackOverflow. One of the many problems that quora face is the duplication of questions. Duplication of question ruins the experience for both the questioner and the answerer. Since the questioner is asking a duplicate question, we can just show him/her the answers to the previous question. And the answerer doesn’t have to repeat his/her answer for essentially the same questions. For example, we have a question like “How can I be a good geologist?” and there are some answers to that question. Later someone else asks another question like “What should I do to be a great geologist?”. We can see that both the questions are asking the same thing. Even though the wordings for the question are different, the intention of both questions is the same. So the answers will be the same for both questions. That means we can just show the answers to the first question. That way the person who is asking the question will get the answers immediately and people who have answered already the first question don’t have to repeat themselves. This problem is available on Kaggle as a competition. https://www.kaggle.com/c/quora-question-pairs So given two questions, our main objective is to find whether they are similar. So let’s do some magic with ML. ? Business Objectives and Constraints There is no strict latency requirement. We would like to have interpretability but it is not absolutely mandatory. The cost of misclassification is medium. Both classes (duplicate or not) are equally important. Data Overview Available Columns: id, qid1, qid2, question1, question2, is_duplicate Class labels: 0, 1 Total training data / No. of rows: 404290 No. of columns: 6 is_duplicate is the dependent variable. No. of non-duplicate data points is 255027 No. of duplicate data points is 149263 We have 404290 training data points. And only 36.92% are positive. That means it is an imbalanced dataset. Business Metrics It is a binary classification. We need to minimize the log loss for this challenge. Basic EDA Test data don’t have question ids. So the independent variables are question1, question2 and the dependent variable is is_duplicate. 3 rows had null values. So We removed them and now We have 404287 question pairs for training. 36.92% of question pairs are duplicates and 63.08% of questions pair non-duplicate. Out of 808574 total questions (including both question1 and question2), 537929 are unique. Most of the questions are repeated very few times. Only a few of them are repeated multiple times. One question is repeated 157 times which is the max number of repetitions. There are some questions with very few characters, which does not make sense. It will be taken care of later with Data Cleaning. Data Cleaning We have converted everything to lower case. We have removed contractions. We have replaced currency symbols with currency names. We have also removed hyperlinks. We have removed non-alphanumeric characters. We have removed inflections with word lemmatizer. We have also removed HTML tags. Feature Extraction We have created 23 features from the questions. We have created features q1_char_num, q2_char_num with count of characters for both questions. We have created features q1_word_num, q2_word_num with count of characters for both questions. We have created total_word_num feature which is equal to sum of q1_word_num and q2_word_num. We have created differ_word_num feature which is absolute difference between q1_word_num and q2_word_num. We have created same_first_word feature which is 1 if both questions have same first word otherwise 0. We have created same_last_word feature which is 1 if both questions have same last word otherwise 0. We have created total_unique_word_num feature which is equal to total number of unique words in both questions. We have created total_unique_word_withoutstopword_num feature which is equal to total number of unique words in both questions without the stop words. The total_unique_word_num_ratio is equal to total_unique_word_num divided by total_word_num. We have created common_word_num feature which is count of total common words in both questions. The common_word_ratio feature is equal to common_word_num divided by total_unique_word_num. The common_word_ratio_min is equal to common_word_num divided by minimum number of words between question 1 and question 2. The common_word_ratio_max is equal to common_word_num divided by maximum number of words between question 1 and question 2. We have created common_word_withoutstopword_num feature which is count of total common words in both questions excluding the stopwords. The common_word_withoutstopword_ratio feature is equal to common_word_withoutstopword_num divided by total_unique_word_withoutstopword_num. The common_word_withoutstopword_ratio_min is equal to common_word_withoutstopword_num divided by minimum number of words between question 1 and question 2 excluding the stopwords. The common_word_withoutstopword_ratio_max is equal to common_word_withoutstopword_num divided by maximum number of words between question 1 and question 2 excluding the stopwords. Then we have extracted fuzz_ratio, fuzz_partial_ratio, fuzz_token_set_ratio and fuzz_token_sort_ratio features with fuzzywuzzy string matching tool. Reference: https://github.com/seatgeek/fuzzywuzzy EDA with Features If First word or Last word is the same then there is a high chance that the question pairs are duplicates. The number of total unique words (q1 and q2 both combined) with and without stopwords is less if question pairs are duplicate. For duplicate question pairs, the total unique words to total words ratio is generally smaller. Duplicate question pairs tend to have more common words between both the questions. Hence extracted features related to common words are also showing differences in distributions. The fuzz ratios tend to be generally higher for duplicate question pairs. Featurization with SentenceBERT We need to convert the questions to some numeric form to apply machine learning models. There are various options from basic like Bag of Words to Universal Sentence Encoder. I tried InferSent sentence embeddings. But it returns 4096 dimension representation. And after applying it the train data became huge. So I discarded it. And I chose SentenceBERT for this problem. SentenceBERT is a BERT based sentence embedding technique. We will use pre-trained SentenceBERT model paraphrase-mpnet-base-v2, which is recommended for best quality. The SentenceBERT produces an output of 768 dimensions. https://www.sbert.net/ We created two more features cosine_simlarity_bert and euclidean_distance_bert which measures similarity and distance between both pairs of questions with SentenceBert representation. The total number of features till now is 25. EDA on new features related toSentenceBERT Cosine Similarity is larger for duplicate pairs. 80% of non-duplicate question pairs and only 20% of duplicate question pairs have cosine similarity of <= .815 Euclidean Distance is smaller for duplicate pairs. 20% of non-duplicate question pairs and approx 80% of duplicate question pairs have euclidean distance of <= 2. It is showing the Pareto Principle (80–20 rule). Data Pre-processing We normalized (min-max scaling) the extracted features. We have not normalized the embeddings because it is not recommended. We have 1561 features (25 + 768 + 768). 25 are extracted features. 768+768 for sentence embedding of question 1 and question 2. Since the dataset was imbalanced. We did oversample by sampling from the minority class. Now we have 510048 data points for training. 255024 from each class. Note that I have not set aside any data for testing locally. Because our main goal is to get a good score on Kaggle. Training Models Support Vector Classifier While training Halving Grid Search CV with param grid, We have used LinearSVC because it is recommended for large datasets. We have used the L2 penalty and the loss function is squared of hinge loss. Also, it is recommended to use primal formulation for large datasets. For some values of C it was not conversing so I increased max_iter to 3000. For cross-validation in halving grid search cv, I have used 1 shuffle split with a 70:30 split. Also, the scoring for selection is accuracy. The halving grid search cv found C=100 to be the best param. And the best accuracy is 85.79%. So the best estimator looks like, Now since we need to minimize log loss for the competition. We would want a good predicted probability. Calibrated Classifier can be used to get a good predicted probability. After calibration of the model for probabilities. I predicted probabilities of test data and submitted on Kaggle. The public leader board score for the Kaggle submission is 0.36980. It is very good considering that the model assumes linear separability. Random Forest You know Quora itself usage Random Forest for this problem. Or at least they did when they first posted the competition on Kaggle in June 2017. Same as before we are using halving grid search cv with following param grid, And the rest of the params are the default for the Random Forest Classifier. We have used the very similar halving grid search cv as before, The halving grid search cv found {‘max_depth’: 150, ‘min_samples_split’: 5, ‘n_estimators’: 800} to be the best params. And the best accuracy is 90.53%. So the accuracy has increased by 5% as compared to SVM. The best estimator looks like, Now at this point, I should have used calibration but because it has already taken a lot of time I skipped it. I should have used Bayesian Optimisation technique ?. The public leader board score for the Kaggle submission is 0.32372, which slightly better than SVC. I was expecting a little less logloss but remember we have not done calibration (due to time constraints). We will try to better with XGBoost — the holy grail of ml models for the Kaggle competition. XGBoost Due to time and system configuration constrained, I decided to use 200000 data points to estimate a few of the params. At first, I was using Optuna for hyperparameter tuning but it had some issues because of which it was not releasing memory after the trials. So the system was crash after few trials. Later on, I decided to use HyperOpt for the tuning. With HyperOpt, I tuned only max_depth and learning_rate. It was not a fine-tune because I used only 5 trials. But it gave a rough idea. Finally, I choose the following params for training the model on whole data, The objective = “binary:logistic” because we are trying to get probabilities. I have used tree_method = “hist” for faster training. grow_policy = “lossguide” is inspired from LightGBM for better accuracy. The num_boost_round is set to 600 with early_stopping_rounds as 20. The public leader board score for the Kaggle submission is 0.32105, which slightly better than the other models. I was expecting a better result than this. Which is possible with more fine-tuning the hyperparameters. XGBoost have tons of hyperparameters https://xgboost.readthedocs.io/en/latest/parameter.html Another XGBoost I was not happy with the result of the XGBoost model so I decided to tune the parameters with gut feeling. The first thing I did is that I got rid of oversampled data by removing the duplicate rows. This time I added a few more parameters to generalize better, Also, I decreased the number of boosting round to 500. ? Voila! We have a winner. This submission resulted in public LB score of 0.28170. This seems a very good result. Final Thoughts I learned a lot from this case study. I took some shortcuts either because of system configuration constraints or some time constraints. I also experienced firsthand that machine learning is not all about model building but steps before that take more time. The hyperparameter tuning can be automated but things like feature extraction or deciding on what featurization to use need to be done manually. I spent almost two weeks ? and half of that time I was waiting for some execution to complete. So I think it’s a good idea to use things like Amazon SageMaker if you have resource-intensive tasks. In the future, we can try some deep learning-based models. References a. https://appliedroots.com/

2026-08-14 李宇轩 人工智能 英-中

Efficient Data Valuation with Exact Shapley Values Produce Better Models with Less Data. This post is an overview of the work done in the paper — ‘Efficient task-specific data valuation for nearest neighbour algorithms’. Does more data always produce better results? Given the trends in machine learning, you would expect as much. More data seems king given the recent rise in larger and larger machine learning models with exceptional performance. This pattern is exemplified in advances in Machine learning like OpenAIs GPT-3 model, trained on over 45 terabytes of data from books and the internet to support its 175 billion parameters. But this model is effectively dwarfed by Google’s new switch Transformer architecture with 1.6 trillion parameters. So it appears that bigger and bigger models with more and more data are inevitable. If data is so valuable, should we aim to gather more? Why shouldn’t we have more of what is so valuable? There is a problem with this approach. What if the additional data you gather is irrelevant to your problem. What if it is purely noise? Then any additional data will actively hurt the performance of your model. So what can you do to mitigate this problem? Well, for one, you can painstakingly look through your data. This approach is probably one of the most effective tools but is highly time-consuming when you have a lot of data. For computer vision tasks, there are a lot of different techniques that you can use. This post shows some exciting techniques for precisely that. For many cases, the applicable data is relatively easy to identify. However, for others, it is an unclear task at best. Often part of the problem is that the relationships between the data and the outcome are what we’re trying to model. Then identifying what data is most applicable is an arduous task. Fortunately, there is a new method that is perfect for this scenario based on Shapley values. SHAP Values SHapley Additive exPlanations (SHAP) is a game-theoretic approach to explain the output of any machine learning model. This method is fairly well known, but the attribution is based solely on the features. The method provides attribution values from features to the output of the model. SHAP values are based on the Shapley values, which determine how to distribute payment among multiple players in a game fairly. For SHAP values, the coalitions of players are based on the features. SHAP values are incredibly flexible. For example, in computer vision tasks, SHAP values represent the attribution of different pixels to the model’s output. There are many different methods to calculate SHAP values, including a KernelSHAP method which is model-independent. For each variation of SHAP values, the attributions are always related to the features of the models. However, there is an alternative. Use the instances as the players and calculate attribution for each instance. Data Valuation ‘Efficient task-specific data valuation for nearest neighbour algorithms’ is a recent paper providing novel algorithms to calculate exact Shapley values. For the rest of this post, I refer to the Shapley values produced for each instance as Data Shapley Values. Data Shapley Values are a recent innovation that utilizes Shapley values to determine the attribution of different data instances. The motivation of the research was inspired by optimizing records selected from a data market for privacy-preserving machine learning. The data market consists of many different medical records. Therefore, the data buyer chooses a subset of records to purchase from the data market. Because of the data cost, the buyer aim’s to select an optimal subset of patients for their model. The Shapley values in this problem configuration measure the marginal improvements of the utility attributed to each data point average overall possible data subsets. The most significant issue for computing Shapley values is the high degree of complexity. Generally, this is on the level of O(2^N) for exact calculations. However, the researchers have developed a novel algorithm for exact Shapley values designed for K-NN classifiers using a KNN utility. This algorithm relies on the fact that the KNN utility satisfies a piecewise utility difference property. I’ll leave the exact mathematical formulation for the paper and the curious readers. But here is the result. The algorithm performs exact computation with O(N log N) complexity. Experiments The structure of the experiments follows a simple format. First, the data is separated into training, validation, and test sets. Then, the exact Shapley values are calculated based on the attribution of the training instances to the validation instances. Then the training instances are ordered according to their Shapley values. This setup provides the user with the attributions of each instance to the validation set. Next, the user can select only those instances with the highest Shapley value, providing the user with a smaller subset of data. As this process selects data that contributes the most to the performance on the validation set, removing instances will low Shapley values often removes the noisiest instances within the data. And, at the same time, maintaining the most representative examples. The experiments use the diamonds dataset. This dataset contains almost 54,000 diamonds. The features include diamond attributes such as the carat of the diamond, the cut, the colour, and several other features. Some of the features are categorical, and for these experiments, I transform these features into boolean features for each category. The target is the price of the diamonds. I’ve taken 5000 instances for validation and testing. The remainder of the data is used for training. The aim is to produce a model to predict the price of diamonds using less data. I’ve evaluated the performance with a single decision tree regressor. The data is ordered based on the Shapley value, and the performance is measured with an R² score. Next, the Shapley values are calculated for the validation. The chunk of code also sequentially fits a model on smaller amounts of data. The data is considering according to decreasing Shapley Values. Once the model is trained on the subset of training data, the model is evaluated on the test dataset. The experiment results show that with less data, the model performs better on the test data. The peak of the test performance occurs when almost 50% of the training data is removed. What is also intriguing is that the model performance on the validation set peaks when over 60% of the training data is removed. Since the instances are removed based on the Shapley values on the validation set, this pattern makes some sense. Conclusion Despite the increasing availability of a massive amount of data, the quality of the data is crucial. Building your models on low-quality data produces low-quality models. Utilizing Shapley values on instances offer an alternative approach. You can create your models with fewer instances and get improved performance. The experiments show that less data can improve the model performance even with a dramatically smaller subset of data. Data is the new oil, but the quality of that oil matters. Consider using data Shapley values in your next machine learning model.

2026-08-12 李宇轩 人工智能 英-中

Explaining how I reached the top ranks of the new Data-Centric competition Following the endorsement of Andrew NG himself(!) regarding my last article, it felt natural to share all tips (with code!) of how I handled DeepLearning.ai's new challenge. The competition explained… again! If you are not familiar yet with the new Data-Centric challenge launched by DeepLearning.ai a few weeks ago, you might have a look at the article I wrote a few weeks ago to describe this challenge. Not sure this is worth it? Just follow Andrew Ng advice ? : And if you are in a hurry, here is the long story short: the objective of the competition is to produce the best possible set of pictures to train a predefined model (ResNet50) to recognize roman numerals. The competition offers a “starting base” of approx. 3000 pictures, including noisy and mislabeled numbers, as you can observe below: Let’s the game begin! I am going to present the different steps to reach a good performance in a smooth way but this is obviously the result of many tests and trials I undertook to find the optimum combination! I have also created a dedicated repository on GitHub (link at the end of the article) if you want to explore my solution further. 1. Pictures review The first task is probably the most demanding: reviewing each picture to check a few criteria. Here are the ones I had used: Does this look like a roman number? (If not, we should remove it!) Is the picture correctly labeled? (ex. “II” in “III” folder or vice versa) What is the number quality? (rating from 1: good to 4:poor) What is the background quality? (same rating as above) What is the font style? (“Arial” or “Roman”) What is the exact format of the number? (“viii” or “VIII”) Could we apply symmetries? (horizontal or vertical symmetries are usually suiting “I, II, III, or X” numbers but not for “i”, “ii”, or “VII”) Here are three examples of my evaluations (stored in a tabular way): While reviewing the 3000 pictures (which took me approximately 2 or 3 hours ?), I sometimes had the feeling of “déjà vu” and I started to wonder whether some duplicates were hidden in the dataset? That would have not been very surprising so I had to take this also into consideration. I also designed a simple function to automatically check the content of the folders to evaluate the results of the different operations I would perform. Like all other functions I would use afterward in the notebook, I stored it in a dedicated “dcc_functions.py” (available on the GitHub repository). Here is the output on the initial dataset: 2. Dataset cleaning 2.1 Noise removal I started by removing all the pictures that I had identified as pure noise or, at least, too noisy to train properly the model. This is obviously a personal choice and each participant has probably ended up with a different selection. I identified approx. 260 pictures to be removed (the corresponding list is stored in an Excel file on the GitHub repo). 2.2 Duplicates removal As explained before, I had the feeling that some pictures were exactly the same but manually identifying them was impossible. There were a few technics I knew that could allow solving this issue: Pairing files with identical sizes… but a lot of false positives would arise Pairing files with identical sizes & configurations (like two “II” or “viii”) Pairing files according to their statistics (using, for ex., PIL’s ImageStat) Pairing files using Structure Similarity Index (some explanations here) Pairing files according to their “hash” number As it was not a “life or death” matter, I decided to use the second solution which was both easy and quick to implement. The script (available here) identified approx 200 pairs of twin pictures, out of which 53 were actually genuine duplicates (some examples below): 2.3 Moving some pictures in the right folders No need to spend a lot of time on that one: when a picture was mislabeled, I simply moved it back to the folder it belongs to. 2.4 Edgy or not edgy? Before we go further, I’d like to share an interesting finding: I reviewed the pictures twice: when I entered the competition and, a second time, when I had a better idea of what to look for in the pictures. When reviewing the original pictures for the first time, I had excluded a lot of “edgy cases” that seemed too ambiguous to train the model. But a few weeks after the competition started, I started to get used to these edgy cases and consider them differently, like: “Well, it could be good to include this one to teach to the model that this case might happen.” I ended up adding approx. 80 pictures to the dataset. Counter-intuitively, the performance was decreasing with this new selection, including more edgy pictures. How come? One of the participants, Mohamed Mohey, highlighted on the dedicated Discourse thread that the 32x32 transformation (applied to the dataset before the training) would sometimes completely denature the essence of the picture, as shown in the example below: We can observe that, due to this 32x32 transformation, an obvious “III” is becoming a plausible “II”, explaining why some edgy cases would not necessarily bring valuable information to the model. It would probably have been a good thing to review the pictures after a 32x32 transformation but I did not! 2.5 Using the “label book” pictures to train the model The organizers from DeepLearning.ai had provided a set of 52 pictures, not existing in the “train” or “validation” folders, to evaluate our model’s performance when the ResNet50 training was over. It was a good way to have a sense of how the model would be performing on the final and hidden dataset but I had also “guessed”, thanks to the scores displayed on the leaderboard, that the final evaluation on the hidden dataset was including 2420 pictures (see the corresponding notebook here). So 52 pictures were not very representative anyway! So I simply included these pictures in my training folder! The merrier, the funnier ? 2.6 Evaluating the impact of the augmentation technics As you might know, it is quite common to use augmentation technics on a dataset composed of pictures to help deep learning models identify the features that allow to properly infer the classes. I decided to consider a few of them: Horizontal and Vertical Symmetries Clockwise and Anti-clockwise rotations (10° and 20°) Horizontal and Vertical Translations Cropping the white areas in the pictures Adding synthetic “salt and pepper” noise Transfering noise of some pictures to some others 2.7 Implementing the customs functions The first functions are quite simple and easily implemented with PIL, OpenCV, or even “packaged solutions” such as ImgAug. I thought it would be more interesting to share some tips regarding some of the custom functions I had designed ? 2.7.1 Squared Cropping Function The cropping operation is an interesting one! As the picture will, ultimately, be converted to a 32x32 picture, it might be better to zoom in on the area where the number is located. However, if the number does not have a “squared” shape, the result could be distorted when converted to 32x32 (as shown below). I redesigned the function so that the cropped output will always have a square shape and avoid this distortion effect: 2.7.2 “Salt and Pepper” Function As the background is probably not always plain white on the final evaluation dataset, I tried to augment pictures by adding a synthetic background. I used the “salt & pepper” function which is, basically, adding “0” and “1” randomly into the NumPy arrays describing the pictures: 2.7.3 Background Noise Transfer Function I was not fully happy with the results of the “Salt and Pepper” function as the noise was always homogeneous so I imagined another way to add noise to the pictures. I recycled some of the pictures that I had originally considered as unreadable and made them become some “noisy background” basis. There were also some pictures with a “heavy background” for which I removed the number (as shown below) to get more samples. It provided me a “noisy backgrounds bank” of 10 pictures that I added randomly to some pictures after applying horizontal or vertical symmetries: 2.8 Choosing the best augmentations As the number of pictures allowed could not exceed 10.000 elements, I had to know which transformations were providing the highest impact. I decided to benchmark them by comparing a baseline (a cleaned dataset with no transformation) and the individual performance of each of the augmentation technics (summary below): We can observe that the rotations, translations, and cropping were bringing a significant impact compared to others so I decided to focus on that ones. And “voilà”! As the process is stochastic (transformations are applied with a 50% probability and some random parameters), each iteration of the script will produce a unique combination of pictures. Many of my tests produced a performance of around 84% while the highest competitor had reached 86% (with a 64% baseline). Honorable I guess ? There would have been some additional tweaks to consider (like creating my own pictures and adding them to the dataset but I choose to only rely on the initial pictures provided). Some others probably gave it a try! A global overview of the competitors’ performance It is probably also worth mentioning that I have analyzed the performance of competitors during the first weeks of the challenge (until 26/08) and we can see how quickly most of the participants reached an acceptable performance, converging quickly towards 75% and above: Final words As mentioned earlier, after a proper review/cleaning of the data and a script executed in less than 30 seconds, you can easily outperform what state-of-the-art models could produce with noisy data! I really enjoyed participating in this challenge which, according to me, was more demanding of “fresh ideas” than “GPU power” and I am really looking forward to the next one! We had a lot of fun and rich interactions with other competitors and Lynn (from DeepLearning.ai) to share our views on this “first-of-its-kind” contest. Many participants were more seeking to share their views and findings along the way rather than being at the top of the leaderboard. I also know how difficult generating “average and noisy data” can be… so congratulations to the organizers for delivering such good material to work on! And, of course, I hope you liked the second part of this Deep-Dive on the Data-Centric Challenge from DeepLearning.ai! As promised, here is the link to the GitHub repository and feel free to share your experience and/or findings in the comments:

2026-08-11 李宇轩 人工智能 英-中

Notes on speech processing, 2.9.2021 Information bottlenecks and dimensionality reduction in deep learning Autoencoders and other deep neural networks with information bottlenecks have become fashionable. The heuristic idea is that the dimensionality of the hidden layers is reduced such that the network is forced to focus on the important part of the data. Experiments have also demonstrated that autoencoders are efficient in this sense. I have however been left wondering whether the amount of information can be characterized in exact terms. How much information flows through the bottleneck? How would we even measure that? This short note is my attempt at characterizing and understanding the problem. I will start with some classical concepts of information theory and linear algebra and then discuss the extent to which such concepts are applicable in machine learning. A central result is that dimensionality of a hidden layer cannot alone be used as a measure of information content. Information content in discrete representations If a system has two states, A and B, then obviously we can represent the state by one bit. Four states can be represented by 2 bits, 8 states by 3 bits and in general, N states by log2(N) bits. We can therefore always easily determine the number of bits required for systems with a finite number of states. The amount of bits needed to describe the state is then a direct measure of the information content or entropy of the system. We can expand this to countable sets, such as integers, if we in additional have the access to probability of each state. Then we can make statements about average bitrate, that is, if we observe the system many times, how many bits do we on average need for representing the state? If the probability of state k is Pk, then the amount of bits needed to represent that state is log2 Pk. That bitrate, log2 Pk occurs at probability Pk, such that the average bitrate can be calculated as the sum, sum Pk log2 Pk, where the summation goes over all k. This applies also when k goes over an infinite but countable set. Information content in linear, continuous valued systems If the title is confusing, just think of linear algebra. How much information is there in a vector x of length N. Well, it is not really defined. What we do however know is that if we multiply it with a matrix A, as y=Ax, then if the matrix A is full rank, then all information is retained. In fact, then we can recover x from y by the inverse x=inv(A)y. No information is lost. Clearly the rank of A thus defines its capacity remove information. If rank(A)<N then information is lost and cannot be recovered from y. This is not yet the whole story though. In practical implementations of the inverse, we know that it is not only the rank which is important, but also the conditioning of A. If any of the singular values of A are close to zero, then A becomes ill-conditioned such that the recovery of x from y becomes numerically difficult. In the best case, we loose accuracy, such that x can be recovered only approximately, in severe cases information can be entirely lost. The information content is thus not only described by dimensionality, but also characterized by accuracy. As we shall see, I argue that it is more useful to characterize loss of information as a loss of accuracy rather than loss of dimensions. Diversion: Space filling curves If you have not heard about space-filling curves, start by watching the Numberphile video about them. The idea is an infinite recursion; you start with a simple shape which goes through a space. Then you add wiggles to that shape so that it spreads more over the space. Repeatedly adding more wiggles makes the curve spread out more and more, such that it converges to covering the whole space. The one-dimensional line thus covers the whole two-dimensional space (i.e. its Hausdorff dimension is 2). In terms of information content, now, the one-dimensional curve contains the information of the two-dimensional space. If we start with some particular point in 2D-space (x,y), we can convert that to a point d on the one-dimensional line, and then convert it back to the 2D-point (x,y). It is just that there is an infinite recursion involved, so this is not a practical algorithm. We can however, implement a finite number of recursions to get an approximation. In the example below, I have implemented an Hilbert-curve and plotted the curve for different number of recursions N. We can readily see that for each iteration, the accuracy with which the curve fills space is doubled (error is halved i.e. error energy is 1/4th). By accuracy I refer to the average distance from a random point in 2D space to the closest point on the curve. Each iteration, on the other hand, splits every segment into 4 sub-segments, at a cost of 2 bits. Halving the error thus comes at a cost of 2 bits. This results thus follows results of conventional lossy coding; halving error costs as many bits as we have dimensions. Now we have 2 dimensions so halving error costs 2 bits. Information content in autoencoders Observe that the above space-filling curve construction can be interpreted as an autoencoder. The 2-dimensional space is mapped (encoder) to a 1 dimensional space (bottleneck), which we can recover with the inverse (decoder). The curve is piecewise linear and could easily be implemented with a single layer of rectified linear units (RELUs). Each recursion consists of a subdivision into 4 parts, such that we can expect that the network can be implemented with 2^(2N) RELUs. Conversely, the error of the mapping is halved if the number of RELUs is quadrupled. A red herring One could easily be fooled to think that we can do some simpler space filling curve than Hilbert (or other equivalent curves). For example, we could draw zig-zag lines going end-to-end on dimension x and then takes a step 1/N on dimension y. This can be implemented with O(N) RELUs. The accuracy of this map would then be relative to 2^-N instead of 2^-(2N). However, we would then have error only on the y dimension and the x dimension could be always perfectly reconstructed. Our accuracy argument thus applies as before, we need 1 bit for each dimension to halve accuracy, when assuming that accuracy on each axis is equal. Reconstruction accuracy as a measure of information The pertinent consequence for autoencoders is that the dimensionality of the bottleneck does not alone define the amount of information that passes through. By exponentially increasing the number of non-linearities in the encoder and decoder, we gain a log-linear decrease in mean square error. Since we thus cannot measure information with the number of dimensions, we should therefore rather measure the amount of information in terms of reconstruction accuracy. This approach is in line also with conventional concepts in probability and statistics. For continuous valued variables x, we cannot define a probability, but only probability distributions, since there are an infinite number of possible values and any particular value would always have probability zero. In a similar fashion, for continuous-valued information bottlenecks, we cannot define absolute information content, but only relative information content, in terms of accuracy. That is, we can say that accuracy (and thus information content) is improved or reduced when changing the network structure, in particular with respect to the number of non-linearities. We can however not say how much information is passed through, but only compare relative amounts of information with different network structures. Vector quantization A particular form of autoencoders which have become fashionable is the VQ-VAE, or vector quantized variational autoencoder. I won’t be going into the ‘variational’ part here, but the vector quantized autoencoder refers to systems where the bottleneck is also quantized. In particular, vector quantizers have a fixed number of quantization levels such that the bitrate is well-defined. The above analysis is thus not directly applicable to such systems. Heuristically, I would argue (and guess) that the encoder complexity has to be sufficient, such that it can digest information into a form which the VQ can handle. Increasing the encoder complexity further would not improve reconstruction accuracy, since it is limited by the VQ accuracy. Conversely, if the encoder has a given structure, then the VQ bitrate has to be sufficient such that it can take full benefit of the embedding. From the space-filling curves above, you can appreciate that if the VQ bitrate is low, then it cannot model the complicated information contained in the high-recursion curves. In other words, the encoder structure and the VQ bitrate have to be jointly matched for optimal performance. Conclusion and to-do’s This is was my first, quick-and-dirty attempt of characterizing the information content in autoencoders. My own impression is that I’m on to something. Clearly a complex encoder can compress information into a narrow bottleneck such that it can be reconstructed with high accuracy. In fact, assuming perfect accuracy (no numerical round-off errors), then any vector could be compressed to a single real value and reconstructed with arbitrary accuracy, if the corresponding encoder and decoder are sufficiently complex. The magic is in the way the space-filling curve embeds infinities; two infinitely accurate signals can be interleaved together without loss of information. The above presentation does not have rigorous proofs and there’s plenty of hand-waving involved. For example, I detailed only the case where a 2D signal is mapped to a 1D signal (2D-to-1D), it can be easily extended to ND-to-1D, but a bit more reflection is needed to extend it to arbitrary width bottlenecks, ND-to-KD. I also did not properly define reconstruction accuracy, nor the number of RELUs in a space-filling curve and so on. I further would like to actually implement the space filling curve with something like pytorch as a demonstration. The VQ discussion was also superficial. I also haven’t done a literature study; let me know if you know of related work! Perhaps next time. In any case, this is a start for a theoretical discussion about information content in autoencoders and related deep neural networks.

2026-08-11 李宇轩 人工智能 英-中

What does word2vec actually learn? And how to train embeddings from similarity functions Representing discrete objects by continuous vectors, the so-called embeddings, has been at the heart of many successful machine learning solutions. The superiority comes from the fact that, unlike the original discrete objects, the embedding vectors offer a compact representation that captures the similarity between the original objects. In this article, we consider the famous word2vec algorithm. Word2vec is simple and intuitive. At a high level, it says that words that appear frequently close to each other should have a similar vector representation. In particular, the example embedding(man) - embedding(king) ~ embedding(woman) - embedding(queen) has become the poster child for the ability of word embeddings to capture word semantics. However, the optimization objective cannot be presented by a well-defined quantity. For comparison, consider learning word embeddings using matrix factorization. Let D be a text corpus consisting of m documents and a vocabulary of n unique words. We compute the n-times-m word, document matrix M, where M[u,v] records how many times word u occurs in document v, see Figure 1. The matrix factorization is defined as In the following, we will slightly abuse notation and denote by u both a word u and its embedding vector. In this case, we know that for a word embedding u, and a document embedding v the inner product between u and v preserves the information how many times the word u occurs in the document v. The larger the embedding dimensionality d, the better the approximation. Unfortunately, there is no such clear formulation of the optimization objective for the word2vec model. What exactly does the inner product of two word vectors in word2vec preserve? And do the embeddings necessarily become better by increasing the dimensionality d? A research paper by Levy and Goldberg answers exactly this question [1]. In this article I present the theoretical results from [1] and later show how they can be used to design a more general class of embeddings. Training word embeddings: word2vec Let us briefly consider how word2vec with negative sampling works. For a more comprehensive description, we refer to this article. Let D be the corpus consisting of word, context pairs. In word2vec the context of word w is defined as the k words surrounding w where k is usually a small constant varying between 5 and 15. We want to learn word embeddings such that if two words frequently co-occur in the corpus, their inner product is large. Consider a word w and let c be a word in its context. For word pairs (w,c) occurring together in the corpus, we want the inner product of the embeddings to maximize the probability that (w,c) indeed appears in the corpus (denoted as D=1). The probability is modeled by a sigmoid function: The above has a trivial solution, we can simply make all inner products arbitrary large. Thus, we also introduce negative pairs, i.e. pairs that do not co-occur in the corpus, for which the objective is: The algorithm can be summarized as follows: We run the above algorithm for several epochs over the corpus in order to guarantee that the learning process converges in an optimum. The theoretical analysis Fix a word w and consider the objective for all pairs in which w appears. Let #(w,c) be the number of appearances of the pair (w,c) in the corpus. We can write the objective as where the second product is over the negative pairs we generate. By taking the logarithm of the objective and observing that each negative word cN has a chance to be sampled, we obtain: Let us explain the above. The word w is fixed, we consider all word context pairs (w,c) that appear in the corpus and we sample k negative pairs (w, cN) such that each word c is sampled with probability #(c)/|D|. We want for positive pairs the inner product to be a large positive number. For negative pairs we want the inner product to be a negative number with large absolute value. Observe that |D|, the number of pairs in the corpus, is constant. Thus, by dividing the above expression by |D| the objective becomes This already provides us with some intuition. The objective is to optimize the embeddings such that they reflect the probability for a positive pair to be sampled as opposed to a pair being sampled at random. Positive pairs are generated with probability And for negative pairs, two words are sampled independently at random, each with probability By setting the inner product as an unknown parameter and solving the corresponding optimization problem, we can find the optimal value for the inner product: In the above P(w,c) is the probability of occurrence of the pair (w,c), and P(w) is the marginal probability of occurrence of word w in the corpus. The above turns out to be a widely used word association measure in natural language processing, the pointwise-mutual information (PMI) measure. This is pretty amazing! It turns out that word2vec is essentially equivalent to matrix factorization where the matrix entries are the PMI scores between word pairs. And PMI as a distance measure was used for NLP-related tasks since the 80s [2], long before the emergence of the concept of word embeddings. Embeddings based on arbitrary similarity functions Now it is easy to see that we can simply replace the probability for sampling positive and negative pairs. We only need to update the second and third steps in the word2vec algorithm presented above: Why is this helpful? This gives us more freedom to assign importance to pairs. We can become creative and consider different similarity measures. For example, the Jaccard similarity between words is defined as follows: Thus, we can learn embeddings that optimize the objective that words w and c are similar to each other if the presence of w implies that it is likely that c also appears in the document, and vice versa. In this case, the pair (“keira”, “knightley”) will likely have a higher score than (“data”, “science”). The objective becomes: And we can also model the probability for generating negative pairs. For example, Pr(w) can be the uniform distribution where all words have the same probability of being selected, disregarding how often they appear. Sampling from a distribution If we could compute and store the similarity for all pairs (u, v), then sampling according to the similarity becomes trivial: just store the pairs with their similarity scores as weights and sample using an algorithm like numpy.random.choice. However, this might be computationally infeasible. There are different approaches to deal with the problem with a larger number of pairs. In general, we want to use as positive pairs only those that have a high similarity score. If your similarity measure is based mainly on counts, then a subsample of the data will preserve the most frequent pairs but many infrequent pairs will be filtered out. For example, we can consider only a subset of the documents in a corpus. Frequent word pairs such as (“data”, “science”) will likely survive. But this might not be the case with (“keira”, “knightley”). For each object, consider only the t nearest neighbors. For example, we might use the publicly available implementation from scikit-learn which uses algorithms like kd-trees to speed up similarity search. These algorithms work well for data that is not very high dimensional. Otherwise, one can consider approaches such as Locality-sensitive hashing that will generate similar words. This is especially true for measures like Jaccard similarity. A practical implementation For illustrative purposes, we implemented a simple solution for learning document embeddings from text corpora. The problem is orthogonal to the problem of training word embeddings: we train vector representations for documents based on the words they contain. We consider the IMDB sentiment analysis dataset. The dataset consists of movie reviews by users and each review is labeled with a positive or negative sentiment. After preprocessing the text, we transformed the documents to vectors by using a tf-idf encoding such that each document The parameter min_df says we consider only words that appear in at least 0.1% of the documents. Essentially, this prevents us from using very specific words that might appear just in a couple of documents. For each input vector, find its t nearest neighbors. This can be achieved using an off-the-shelf package such as scikit-learn’s K-NearestNeighbor which returns nearest neighbors for: Compute the similarities for the generated n*t positive pairs, sort them in an array, and sample according to their weight using numpy.random.choice(): Use a Keras generator to generate positive and negative pairs: Feed the generated pairs into a shallow neural network with an embeddings layer, a Dot layer computing the inner product, and an output layer with a sigmoid activation function: The above approach will train embeddings: Then we can extract the embedding layer for each word and cluster the documents (similarly to what is shown in the gif in Figure 1). We observe that the sentiment distribution in the two clusters is very different: Code The Python implementation for the above is publicly available at: https://github.com/konstantinkutzkov/sim2vec [1] Omer Levy, Yoav Goldberg. Neural Word Embedding as Implicit Matrix Factorization. NIPS 2014: 2177-2185 [2] Kenneth Ward Church and Patrick Hanks. Word association norms, mutual information, and lexicography. Computational linguistics, 16(1):22–29, 1990.

2026-08-11 李宇轩 人工智能 英-中

Face Landmark Detection using Python The comparison between dlib and mediapipe library. Introduction Face landmark detection is a computer vision task where we want to detect and track keypoints from a human face. This task applies to many problems. For example, we can use the keypoints for detecting a human’s head pose position and rotation. With that, we can track whether a driver is paying attention or not. Also, we can use the keypoints for applying an augmented reality easier. And there are so many solutions that we can generate based on this task. Thankfully, we don’t have to understand the concepts of face landmark detection in detail. We can use the prebuilt library like dlib, OpenCV, and mediapipe. In this article, I will show you how to implement face landmark detection with dlib and mediapipe. Without further, let’s get started! Face Landmark Detection with Dlib Dlib is a library for applying machine learning and computer vision solutions. This library is based on the C++ language, but we can use a language like Python for using the library. One of the solutions that we can apply by using this library is face landmark detection. Now let’s get into the implementation. Install the library Installing a library can become a problem. If we don’t have a good guide, installing the library can take several days. Dlib is one of them. Because it uses C++ as the primary language, we have to install C++ tools for installing the library. There are several steps that we should do to install it. Here are the steps: First, install the CMake. You can download the software here. If you are using Windows, please locate the CMake file path first. Then, set the path to the executable path on the environment variable. Then, install Visual Studio with the C++ dependencies to it. You can download the software here. For the dependencies, you can look at this screenshot below: After you install the Visual Studio, the next step is to install the Python. To make your installation simpler, I recommend you for installing Anaconda. You can download it here. For the Python version, I recommend you for using the 3.6.6 version to avoid any errors. Lastly, install the CMake, dlib, and OpenCV library by using pip. Here is the command for doing that: Import the libraries After we’ve installed the libraries, the next step is to import them into our code. We will import OpenCV for retrieving inputs from the webcam, NumPy for numerical computation, and Dlib for detecting keypoints from a face. Here is the code for doing that: Initialize the objects Now let’s initialize several variables. There are three must need variables that we will initialize: A detector for detecting one or more faces. We set the dlib.get_frontal_face_detector function inside the variable. A predictor for detecting keypoints from faces. We set the dlib.shape_predictor function inside the variable. This function needs a pretrained model location as the parameter, which you can download here. The cv2.VideoCapture object for capturing images from the webcam. Also, we set a parameter with value 0 for capturing images from a webcam. Let’s write this code for initializing variables: Face landmark detection mechanism As you can see from above, we initialize the face landmark detector by using the pretrained model. The model is based on ensemble regression trees because the model will predict continuous numbers. You can read the details about the model here. That model is trained on the iBUG-300 W dataset, where it contains images and their corresponding 68 face landmark points. In general, those landmark points belong to the nose, the eyes, the mouth, and the edge of a face. You can download the dataset here. Here is the visualization of the face landmark locations below: Implement the face landmark detection Now you know how the face landmark detection algorithm works. Now let’s implement the algorithm. For implementing that, you can see the code below along with explanations on each line of code: By combining all the code as one, now let’s try the code! If the code doesn’t have any errors, the webcam will display the result along with the keypoints. In my case, here is the result: Face Landmark Detection with Mediapipe Mediapipe is a tool for implementing ML-based computer vision solutions. The tool is created by Google. This tool contains varieties computer vision solutions, such as face detection, pose estimation, object detection, and many more. The advantage of this library is that you can apply the solutions on many platforms, such as web, mobile, PC, and many more. I’ve already explained in the previous section to you how to implement face landmark detection using dlib. Now let’s implement the face landmark detection using Mediapipe. The mechanism The library uses the BlazeFace model for detecting face landmarks. BlazeFace is a deep learning model that is already optimized for low spec devices like smartphones. Therefore, we can use the model in real-time. BlazeFace contains two main steps. First, the model detects one or more faces on an image. Second, the image detects around 468 face keypoints by using regression. Different from the dlib library, this model detects 3D coordinates. Those x and y coordinates are normalized from the image scale. The z coordinate is retrieved by taking the relative calculation between the screen and the model x coordinates. You can read more details here. Here is the flattened mesh from a face with their corresponding indexes: Implementing the face landmark detection In general, the pipeline for implementing face landmark detection is the same as the dlib library. It starts from importing libraries, initializing objects, detect face and its landmarks, and done. Here is the code for doing that: If you implement the code correctly, the image will display on your computer. Here is the preview of my result: The comparison We have already take a walkthrough of face landmark detection libraries using dlib and mediapipe. We can say that both libraries are easy to use. Therefore, we can build our solution rapidly. However, there are differences between them. The dlib library needs C++ dependencies it. That’s why we need CMake and Visual Studio for installing the library. Also, this library needs a specific python library. Therefore, you have to create a virtual environment if you don’t have the supported Python version to run the library. On the other side, installing mediapipe is easier. All you need to do is to install from pip only. Therefore, you don’t have to worry about installation more while using the mediapipe. In the case of the solution, the dlib can detect only the 2D coordinates of the keypoints. On the other hand, the mediapipe can detect the 3D coordinates of the keypoints. Therefore, you can use those keypoints from the mediapipe library for estimating the head pose. Final Remarks Well done! Now you know how to implement face landmark detection using Python. I’ve shown you the libraries like dlib and mediapipe for implementing the solution. I hope it helps you in implementing a computer vision solution. Also, I hope it can become your foundation to build more complex applications. If you are interested in my articles, you can follow me on Medium for more articles like this. Also, if you have any questions, you can contact me on LinkedIn. Thank you for reading my article! References [1] https://www.pyimagesearch.com/2017/04/03/facial-landmarks-dlib-opencv-python/ [2] https://www.analyticsvidhya.com/blog/2021/07/facial-landmark-detection-simplified-with-opencv/

2026-08-11 李宇轩 人工智能 英-中

Topic Model Based Recommendation Systems A very quick and (hopefully) easy to follow introduction into the intuition (and very low level Maths) involved in Topic Model Based Recommendation Systems. Check out my GitHub for a working simple recommendation system based on Topic Modelling. In todays world, sometimes it feels like we are plagued with never ending decisions. Whether it be the Friday night movie or the next song to keep people dancing at an NYE party. So how do recommendation systems actually work? In this article I’m going to explain one approach based on Topic Modelling using a Latent Dirichlet Allocation (LDA). Topic Modelling Before we talk about how to model a topic, we need to first understand what a topic actually is. This is not an intuitive idea to think about so we will describe it in terms of collections of words. If we have a collection of documents randomly selected from a database, we can imagine that some of the words contained in these documents may be semantically similar, or be related to the same area. For example, if these documents were a collection of film reviews. We can imagine that we might be able to form groups of positive and negative reviews based on the words contained within them. Alternatively, we may wish to form collections of documents relating to Sci-Fi, Comedy, Romance etc. So, as we are starting to reorganise this collection of documents into many smaller collections. At the same time, we are starting to see that there are many layers of possibilities. Where each possibility is a selection of topics. You may then ask the question, how can we get a computer to organise these documents into topics and how do we know what topics it will pick? To answer this, we are going to think at a slightly deeper level… The Less General Idea In this article, I am going to describe one option for how to get a computer to perform topic modelling. However there are many other algorithms, approaches and methodologies out in the wild. Going back to our collection of documents and sticking with the approach of looking at the words contained within them, we can build up a vocabulary containing all of the unique words in the database. Say we have 1000 different unique words across 4 documents and we want to characterise each document by which words are contained within them (and by how many of each word). So now we can imagine that for each document, we have a vector with dimension 1000 (one dimension for each unique word). And at each position in each vector there is the count of how many times the word that this position corresponds to, appears in the document. For example, the first position in the vector corresponds to the first word in the vocabulary, which we will say is “Robot”. The first document is a film review about Terminator 23 (or whatever number we are on now…) and so the word “Robot” is mentioned 19 times. Therefore in the first position of the vector corresponding to the first document, we have (19,…). The second position corresponds to the word “Sport” and is mentioned zero times in the Terminator review which gives us (19, 0, …) and so on… The second document happens to be a review for a Tennis documentary and so for this document we have the vector (0, 10, …), since “Robot” is mentioned 0 times. Too many topics? Yes, way too many. I agree, as will your computer. At the moment we have essentially defined a “topic” for each word in the vocabulary. Which clearly is not ideal and will not provide us with many clearly separated topics to play with later on. The next step then is to find a middle ground where each of our reviews belong to a broader topic which is defined by a number of words in the vocabulary. Latent Dirichlet Allocation We are now looking to reduce the number of connections coming from each document by introducing a hidden layer of topics between the individual words in the vocabulary. This is exactly what we are going to use Latent Dirichlet Allocation (LDA) for. LDA requires us to define a required number of topics we want, this is what we call a hyperparameter (a parameter which is defined before the algorithm is run). In a real world we can imagine scanning over many numbers of topics to find the best outcome (aptly called a hyperparameter scan). Let’s say we are looking for 10 topics. This means we wish to add a hidden layer between the previous connections we had, which were linking documents to words (remembering our large vectors (19, 0, …) and (0, 10, …)). Mathematically and computationally this is very desirable for us, since we can replace these huge vectors describing each document with new vectors of size 10 (or however many topics you have chosen). To describe what the LDA is attempting to achieve, it is easiest to look at a matrix formulation below… So in our original perfect description of the documents, we had the matrix S. This matrix is the most complete picture we can have of all the documents, there is no information lost since each word is its own topic and if we ignore the orderings of the words, we are able to perfectly recreate each document. However, some of the information is too fine grained. For example, we don’t really need a separate topic for “Robot” and “Android”. These can just be combined into a coarser topic of “Sci-Fi” or whatever you want to name it. In this case, we can see that we have sacrificed some information. So if we were to recreate the document, there is no guarantee that we would get back the word “Robot”, since we only have information that a similar word from the same topic was mentioned. This is what matrices M and N are doing in this case. The Latent dimension K is our hidden layer of topics (i.e. 10 topics). And given this dimension K, the LDA is learning the matrices M and N in an attempt to best recreate the matrix S. It is not essential you understand the maths here to know what is going on, the takeaway points are: LDA creates coarser grained topics based on the documents given to it. As the LDA model does this, we lose specific information about the individual documents. If you think about how you would sort film reviews into 10 topics, this will hopefully start to make sense. Imagine if I asked you to summarise one of the topics you had created, you wouldn’t be able to recite every word of each document, but you’d probably be able to give a few of the most common words describing the overall topic. Hence, you have lost information. Why do we want to lose information? Losing information may sound like a bad thing, but it actually helps Machine Learning models find patterns in data that they would have missed otherwise. The trick is to not lose the useful information, which in some cases can be learnt in an algorithm or it must be controlled via a hyperparameter (as for LDA with the number of topics). Let’s look at an example… say we have a noisy set of data points loosely describing a quadratic function (plot A). In plot B, here we have tried to keep all the information we have to describe the data points, but does this look right? Probably not, our model is too specific and has not really caught the general trend. In plot C, we have lost too much information. The model is too simple for the data. In plot D, we have a good description of the data. We have found a good hyperparameter to be able to lose the less useful information and keep the important parts. This may all sound familiar because it is exactly the description of underfitting and overfitting data, just in the context of NLP and topic modelling. Back to Topics… Hopefully that brief interlude was useful. If not, sorry about that but we’re back on track now! So we have our LDA model which has sorted all these documents into a distribution of topics. Note that these are not hard clustered topics, they are distributions. So if 3 of the 10 topics we had were Sci-Fi, Documentary and Technology, we could have a film review for a RomCom between astronauts that would have a distribution of (0.2, 0.4, 0.3,…). On the diagram above, this corresponds to the top layer of connections between the documents and the topics. These distributions are actually called Embeddings. Since we have embedded information about the document into a usable mathematical format. Take our embedding from earlier, a = (0.2, 0.4, 0.3,…). Now if we have two more documents with embeddings b = (0.1, 0.5, 0.2,…) and c = (0.8, 0, 0.1,…). Is b or c more similar to embedding a? There are a few ways to answer this question but we will choose the simple (and very effective) cosine similarity. Which you may remember from Maths courses at school or college. This basically calculates the distance between the two embeddings and if they are closer together, they are more similar. In this case a and b are more similar than a and c. So if we were to ask the computer which one out of b and c would it recommend given a. We would expect that it may recommend to us the document b. This is how we can use topic modelling to create recommendations. Round up So we have looked at what topics are, then at what the LDA algorithm gives us and then finally how we can use these mathematical objects to produce recommendations. The key point for recommendations in topic modelling based on similarity is that we are assessing how similar the encoded information of each document is. As we have discussed, these recommendations may be completely terrible depending on how our topics are found and distributed (plot B or plot C from the crude explanation earlier). Or they might be fantastic and hit the sweet spot (plot D). The other point to note is that although we can control the number of topics, we have less control on what these topics are. This is determined from the data, which we are able to manipulate by removing unimportant words for example. But if we don’t have a representative sample of Sci-Fi film reviews in our database then the likelihood is that this topic will not exist. Another reason why its all about the DATA… This is a really short and low level insight into how these types of algorithms can be used to give recommendations. There is loads of much more extensive descriptions out there so I encourage you to read around. If you are interested, I have a working recommendation system code on my GitHub. References [1] — D. Blei, et. al., Latent Dirichlet Allocation (2003)

2026-08-11 李宇轩 人工智能 英-中

Churn prediction model Musing about a use case that’s been with me for a decade No company likes to lose valuable customers. In the beginning, a company typically focuses on acquiring new clients, then grows by offering additional products to existing clients or trying to get them to use their products more. If all is going well, there comes a point when the company is large enough that it must also choose a slightly more defensive strategy and focus on retaining existing customers. Despite the best user experience, there will always be a group of clients who are not satisfied and decide to leave. The company then faces the problem of how to prevent these (voluntary) departures as effectively as possible. This is where the churn model, among others, comes to the rescue. What is the churn model? It’s a predictive model that estimates — at the level of individual customers — the propensity (or susceptibility) they have to leave. For each customer at any given time, it tells us how high the risk is of losing them in the future. Technically, it’s a binary classifier that divides clients into two groups (classes) — those who leave and those who don’t. In addition to assigning them to one of the two groups, it will typically give us the probability with which the client belongs to that group. It is important to note that this is the probability of belonging to the group of clients who leave. Thus, it is the propensity to leave and not the probability of leaving. However, it is possible to estimate the probability through a churn model. What is it useful for? By knowing which clients are at the highest risk of leaving, we can better target our rescue efforts. For example, we can reach out to these clients with a marketing campaign, reminding them that they haven’t purchased from us in a while, or even offering them a benefit. In addition to knowing which clients to target, we can use the churn model to calculate the maximum benefit price that is still worthwhile. For example, if we know that the estimated probability of a particular client leaving is 10% and their annual revenue is $100, the expected value of future annual revenue is $90. Therefore, an offer that typically reduces the probability of leaving to 5% (the expected value of the revenue is then $95) will be worthwhile for this client, so long as it does not cost more than $5. What do we need for the churn model? Like any supervised machine learning model, a churn model needs training data with response (target) and explanatory variables (features). Based on this training data, the model learns to best capture the relationship between features and target. Typically, this is historical data, where we know which clients eventually left and which did not. Those who left have a positive target (yes, they left). Others have a negative target (no, they didn’t leave). Whilst features describe clients at a point in time when that outcome was not yet known. A properly defined target is fundamentally key. In many cases this is simple (e.g., cancellation of last product), sometimes less so (e.g., no transactions in the last three months). However, it is possible to apply the churn model to both contractual (e.g., bank) and non-contractual (e.g., e-shop) client relationships. Features include any data that can help identify clients who churn. Often this includes socio-demographic data, data on products owned, historical transactions, client-company interaction, e-commerce behaviour, and so on. It is also important to be careful about how far in advance we want to estimate the propensity to leave. In other words, how long is the time between the day we look at clients through the available features and the day we can tell if they have left? If that time is too short, we won’t have much time to make any kind of response. If, on the other hand, it is too long, the model will be less accurate and up to date. What does such a model look like? Modern churn models are often based on machine learning; specifically, on the binary classification algorithms mentioned above. There are a number of these algorithms, and it is necessary to test which one best fits a specific situation (specific training data, amount of data, etc.). Whether you use simple models such as logistic regression, more complex random forest or GBM, or venture into neural networks, you need to pay attention to the following two things. Classifiers have a variety of performance metrics. Since churn is very low for most companies, it is not enough to look at the accuracy of the churn model. For example, if the churn is 10% and the churn model for all clients says they will not leave, it will have 90% accuracy. But this is not useful. So, among other things, you need to look at sensitivity (how many of the clients who actually leave were detected by the model) and precision (how many of the clients identified by the model actually left). Furthermore, it is advisable to not use the resulting model as a black box. Rather, try to understand the parameters based on which decisions are made. Not only can this reveal flaws in the model or data, but it can also be very useful information for product and marketing teams. For example, if we know that the absolute amount of discount has less impact on churn than the relative amount of discount, we can use this to create more effective campaigns and pricing strategies. What next? Once you have the churn model ready, you need to plug it into the day-to-day running of the company. This involves monitoring, evaluating and updating it on an ongoing basis (whether that’s simply re-training it or even adding new features). Consequently, you can start to automatically detect events that tend to increase the propensity to leave that need to be responded to as quickly as possible. External data consultants can help you with both. But beware, it is crucial for the churn model (even more than for other data projects) to involve people with experience and a feel for the specific situation in the company and industry. The article was originally written in Czech and published on Bizztreat Blog. As ever, I’m indefinitely grateful to Chelsea Wilkinson for patiently shaping my thoughts into a publishable format. Thanks for reading! Please feel free to share your thoughts or opinions in the comments. Follow me on Medium, LinkedIn and Twitter.

2026-08-11 李宇轩 人工智能 英-中

XGBoost For Time Series Forecasting: Don’t Use It Blindly Forecasting techniques don’t work well with all time series When modelling a time series with a model such as ARIMA, we often pay careful attention to factors such as seasonality, trend, the appropriate time periods to use, among other factors. However, when it comes to using a machine learning model such as XGBoost to forecast a time series — all common sense seems to go out the window. Rather, we simply load the data into the model in a black-box like fashion and expect it to magically give us accurate output. A little known secret of time series analysis — not all time series can be forecast, no matter how good the model. Attempting to do so can often lead to spurious or misleading forecasts. To illustrate this point, let us see how XGBoost (specifically XGBRegressor) varies when it comes to forecasting 1) electricity consumption patterns for the Dublin City Council Civic Offices, Ireland and 2) quarterly condo sales for the Manhattan Valley. How XGBRegressor Forecasts Time Series XGBRegressor uses a number of gradient boosted trees (referred to as n_estimators in the model) to predict the value of a dependent variable. This is done through combining decision trees (which individually are weak learners) to form a combined strong learner. When forecasting a time series, the model uses what is known as a lookback period to forecast for a number of steps forward. For instance, if a lookback period of 1 is used, then the X_train (or independent variable) uses lagged values of the time series regressed against the time series at time t (Y_train) in order to forecast future values. Forecasting Electricity Consumption Let’s see how this works using the example of electricity consumption forecasting. The dataset in question is available from data.gov.ie. From this graph, we can see that a possible short-term seasonal factor could be present in the data, given that we are seeing significant fluctuations in consumption trends on a regular basis. Let’s use an autocorrelation function to investigate further. From this autocorrelation function, it is apparent that there is a strong correlation every 7 lags. Intuitively, this makes sense because we would expect that for a commercial building, consumption would peak on a weekday (most likely Monday), with consumption dropping at the weekends. When forecasting such a time series with XGBRegressor, this means that a value of 7 can be used as the lookback period. The model is run on the training data and the predictions are made: Let’s calculate the RMSE and compare it to the test mean (the lower the value of the former compared to the latter, the better). We see that the RMSE is quite low compared to the mean (11% of the size of the mean overall), which means that XGBoost did quite a good job at predicting the values of the test set. If you wish to view this example in more detail, further analysis is available here. Forecasting Manhattan Valley Condo Sales In the above example, we evidently had a weekly seasonal factor, and this meant that an appropriate lookback period could be used to make a forecast. However, there are many time series that do not have a seasonal factor. This makes it more difficult for any type of model to forecast such a time series — the lack of periodic fluctuations in the series causes significant issues in this regard. Here is a visual overview of quarterly condo sales in the Manhattan Valley from 2003 to 2015. The data was sourced from NYC Open Data, and the sale prices for Condos — Elevator Apartments across the Manhattan Valley were aggregated by quarter from 2003 to 2015. From the above, we can see that there are certain quarters where sales tend to reach a peak — but there does not seem to be a regular frequency by which this occurs. Again, let’s look at an autocorrelation function. From the autocorrelation, it looks as though there are small peaks in correlations every 9 lags — but these lie within the shaded region of the autocorrelation function and thus are not statistically significant. What if we tried to forecast quarterly sales using a lookback period of 9 for the XGBRegressor model? The same model as in the previous example is specified: Now, let’s calculate the RMSE and compare it to the mean value calculated across the test set: We can see that in this instance, the RMSE is quite sizable — accounting for 50% of the mean value as calculated across the test set. This indicates that the model does not have much predictive power in forecasting quarterly total sales of Manhattan Valley condos. Given that no seasonality seems to be present, how about if we shorten the lookback period? Let’s try a lookback period of 1, whereby only the immediate previous value is used. The size of the mean across the test set has decreased, since there are now more values included in the test set as a result of a lower lookback period. This has smoothed out the effects of the peaks in sales somewhat. However, we see that the size of the RMSE has not decreased that much, and the size of the error now accounts for over 60% of the total size of the mean. Therefore, using XGBRegressor (even with varying lookback periods) has not done a good job at forecasting non-seasonal data. Conclusion There are many types of time series that are simply too volatile or otherwise not suited to being forecasted outright. However, all too often, machine learning models like XGBoost are treated in a plug-and-play like manner, whereby the data is fed into the model without any consideration as to whether the data itself is suitable for analysis. Therefore, the main takeaway of this article is that whether you are using an XGBoost model — or any model for that matter — ensure that the time series itself is firstly analysed on its own merits. This means determining an overall trend and whether a seasonal pattern is present. The allure of XGBoost is that one can potentially use the model to forecast a time series without having to understand the technical components of that time series — and this is not the case. Many thanks for your time, and any questions or feedback are greatly appreciated. Disclaimer: This article is written on an “as is” basis and without warranty. It was written with the intention of providing an overview of data science concepts, and should not be interpreted as professional advice. The findings and interpretations in this article are those of the author and are not endorsed by or affiliated with any third-party mentioned in this article. The author has no relationship with any third parties mentioned in this article.

2026-08-11 李宇轩 人工智能 英-中

Posted by Katrin Tomanek, Software Engineer and Bob MacDonald, Technical Program Manager, Google Research Speech impairments affect millions of people, with underlying causes ranging from neurological or genetic conditions to physical impairment, brain damage or hearing loss. Similarly, the resulting speech patterns are diverse, including stuttering, dysarthria, apraxia, etc., and can have a detrimental impact on self-expression, participation in society and access to voice-enabled technologies. Automatic speech recognition (ASR) technologies have the potential to help individuals with such speech impairments by improving access to dictation and home automation and by enhancing communication. However, while the increased computational power of deep learning systems and the availability of large training datasets has improved the accuracy of ASR systems, their performance is still insufficient for many people with speech disorders, rendering the technology unusable for many of the speakers who could benefit the most. In 2019, we introduced Project Euphonia and discussed how we could use personalized ASR models of disordered speech to achieve accuracies on par with non-personalized ASR on typical speech. Today we share the results of two studies, presented at Interspeech 2021, that aim to expand the availability of personalized ASR models to more users. In “Disordered Speech Data Collection: Lessons Learned at 1 Million Utterances from Project Euphonia”, we present a greatly expanded collection of disordered speech data, composed of over 1 million utterances. Then, in “Automatic Speech Recognition of Disordered Speech: Personalized models outperforming human listeners on short phrases”, we discuss our efforts to generate personalized ASR models based on this corpus. This approach leads to highly accurate models that can achieve up to 85% improvement to the word error rate (WER) in select domains compared to out-of-the-box speech models trained on typical speech. Impaired Speech Data Collection Since 2019, speakers with speech impairments of varying degrees of severity across a variety of conditions have provided voice samples to support Project Euphonia’s research mission. This effort has grown Euphonia’s corpus to over 1 million utterances, comprising over 1400 hours from 1330 speakers (as of August 2021). Distribution of severity of speech disorder and condition across all speakers with more than 300 utterances recorded. For conditions, only those with > 5 speakers are shown (all others aggregated into “OTHER” for k-anonymity). ALS = amyotrophic lateral sclerosis; DS = Down syndrome; PD = Parkinson’s disease; CP = cerebral palsy; HI = hearing impaired; MD = muscular dystrophy; MS = multiple sclerosis To simplify the data collection, participants used an at-home recording system on their personal hardware (laptop or phone, with and without headphones), instead of an idealized lab-based setting that would collect studio quality recordings. To reduce transcription cost, while still maintaining high transcript conformity, we prioritized scripted speech. Participants read prompts shown on a browser-based recording tool. Phrase prompts covered use-cases like home automation (“Turn on the TV.”), caregiver conversations (“I am hungry.”) and informal conversations (“How are you doing? Did you have a nice day?”). Most participants received a list of 1500 phrases, which included 1100 unique phrases along with 100 phrases that were each repeated four more times. Speech professionals conducted a comprehensive auditory-perceptual speech assessment while listening to a subset of utterances for every speaker providing the following speaker-level metadata: speech disorder type (e.g., stuttering, dysarthria, apraxia), rating of 24 features of abnormal speech (e.g., hypernasality, articulatory imprecision, dysprosody), as well as recording quality assessments of both technical (e.g., signal dropouts, segmentation problems) and acoustic (e.g., environmental noise, secondary speaker crosstalk) features. Personalized ASR Models This expanded impaired speech dataset is the foundation of our new approach to personalized ASR models for disordered speech. Each personalized model uses a standard end-to-end, RNN-Transducer (RNN-T) ASR model that is fine-tuned using data from the target speaker only. Architecture of RNN-Transducer. In our case, the encoder network consists of 8 layers and the predictor network consists of 2 layers of uni-directional LSTM cells. To accomplish this, we focus on adapting the encoder network, i.e. the part of the model dealing with the specific acoustics of a given speaker, as speech sound disorders were most common in our corpus. We found that only updating the bottom five (out of eight) encoder layers while freezing the top three encoder layers (as well as the joint layer and decoder layers) led to the best results and effectively avoided overfitting. To make these models more robust against background noise and other acoustic effects, we employ a configuration of SpecAugment specifically tuned to the prevailing characteristics of disordered speech. Further, we found that the choice of the pre-trained base model was critical. A base model trained on a large and diverse corpus of typical speech (multiple domains and acoustic conditions) proved to work best for our scenario. Results We trained personalized ASR models for ~430 speakers who recorded at least 300 utterances. 10% of utterances were held out as a test set (with no phrase overlap) on which we calculated the word error rate (WER) for the personalized model and the unadapted base model. Overall, our personalization approach yields significant improvements across all severity levels and conditions. Even for severely impaired speech, the median WER for short phrases from the home automation domain dropped from around 89% to 13%. Substantial accuracy improvements were also seen across other domains such as conversational and caregiver. WER of unadapted and personalized ASR models on home automation phrases. To understand when personalization does not work well, we analyzed several subgroups: HighWER and LowWER: Speakers with high and low personalized model WERs based on the 1st and 5th quintiles of the WER distribution. SurpHighWER: Speakers with a surprisingly high WER (participants with typical speech or mild speech impairment of the HighWER group). Different pathologies and speech disorder presentations are expected to impact ASR non-uniformly. The distribution of speech disorder types within the HighWER group indicates that dysarthria due to cerebral palsy was particularly difficult to model. Not surprisingly, median severity was also higher in this group. To identify the speaker-specific and technical factors that impact ASR accuracy, we examined the differences (Cohen's D) in the metadata between the participants that had poor (HighWER) and excellent (LowWER) ASR performance. As expected, overall speech severity was significantly lower in the LowWER group than in the HighWER group (p < 0.01). Intelligibility and severity were the most prominent atypical speech features in the HighWER group; however, other speech features also emerged, including abnormal prosody, articulation, and phonation. These speech features are known to degrade overall speech intelligibility. The SurpHighWER group had fewer training utterances and lower SNR compared with the LowWER group (p < 0.01) resulting in large (negative) effect sizes, with all other factors having small effect sizes, except fastness. In contrast, the HighWER group exhibited medium to large differences across all factors. Speech disorder and technical metadata effect sizes for the HighWER-vs-LowWER and SurpHighWER-vs-LowWER pairs. Positive effects indicated that the group values of the HighWER group were greater than LowWER groups. We then compared personalized ASR models to human listeners. Three speech professionals independently transcribed 30 utterances per speaker. We found that WERs were, on average, lower for personalized ASR models compared to the WERs of human listeners, with gains increasing by severity. Delta between the WERs of the personalized ASR models and the human listeners. Negative values indicate that personalized ASR performs better than human (expert) listeners. Conclusions With over 1 million utterances, Euphonia’s corpus is one of the largest and most diversely disordered speech corpora (in terms of disorder types and severities) and has enabled significant advances in ASR accuracy for these types of atypical speech. Our results demonstrate the efficacy of personalized ASR models for recognizing a wide range of speech impairments and severities, with potential for making ASR available to a wider population of users. Acknowledgements Key contributors to this project include Michael Brenner, Julie Cattiau, Richard Cave, Jordan Green, Rus Heywood, Pan-Pan Jiang, Anton Kast, Marilyn Ladewig, Bob MacDonald, Phil Nelson, Katie Seaver, Jimmy Tobin, and Katrin Tomanek. We gratefully acknowledge the support Project Euphonia received from members of many speech research teams across Google, including Françoise Beaufays, Fadi Biadsy, Dotan Emanuel, Khe Chai Sim, Pedro Moreno Mengibar, Arun Narayanan, Hasim Sak, Suzan Schwartz, Joel Shor, and many others. And most importantly, we wanted to say a huge thank you to the over 1300 participants who recorded speech samples and the many advocacy groups who helped us connect with these participants.

2026-08-09 李宇轩 人工智能 英-中

本校翻译实践排行榜

梁雯
211368字
范荣
146854字
孟焕蕊
122789字
4
李宇轩
34875字
5
王铅铅
17717字

400所高校都在用的翻译教学平台

试译宝所属母公司