【发布时间】:2014-08-18 11:15:14
【问题描述】:
我在 AWS Elastic MapReduce 上运行以下 MapReduce:
./elastic-mapreduce --create --stream --name CLI_FLOW_LARGE --mapper s3://classify.mysite.com/mapper.py --reducer s3://classify.mysite.com/reducer.py --input s3n://classify.mysite.com/s3_list.txt --output s3://classify.mysite.com/dat_output4/ --cache s3n://classify.mysite.com/classifier.py#classifier.py --cache-archive s3n://classify.mysite.com/policies.tar.gz#policies --bootstrap-action s3://classify.mysite.com/bootstrap.sh --enable-debugging --master-instance-type m1.large --slave-instance-type m1.large --instance-type m1.large
由于某种原因,cacheFile classifier.py 似乎没有被缓存。当reducer.py 尝试导入它时出现此错误:
File "/mnt/var/lib/hadoop/mapred/taskTracker/hadoop/jobcache/job_201204290242_0001/attempt_201204290242_0001_r_000000_0/work/./reducer.py", line 12, in <module>
from classifier import text_from_html, train_classifiers
ImportError: No module named classifier
classifier.py 绝对存在于s3n://classify.mysite.com/classifier.py。值得一提的是,政策存档似乎加载得很好。
【问题讨论】:
-
我认为它可能被缓存了,但是你的 python 路径有问题。 reducer.py 和 classifier.py 是否被缓存到同一个目录中?如果没有,您将不得不摆弄
sys.path -
我以为缓存的东西在工作目录中?
-
其实你是对的!发布我现在找到的解决方案..
标签: python hadoop amazon-web-services elastic-map-reduce